Unknown Company

Application Reliability engineer

santa clara, ca • Posted 3 days ago
Onsite Full Time General

Application Reliability EngineerLocation: Sunnyvale, CA. Duration: 6+Essential Job Functions5+ years of overall experience in providing application development and support for on-premise and cloud based Enterprise Data warehouse applications3+ years of prior experience working as a software developerGood programming skills in at least 2 languages Java, Perl, Python, ScalaGood hands-on experience in Unix, Linux, Shell scripting, SQL, autosys, Kafka and SplunkHands-on with development to Production processes including testing, version control tools like git/svn and experience with source control, continuous integration, deployments etc.Continuously identify, elevate and lead the risks and mitigation paths for high-velocity and complex NPIs.Build a deep understanding of multi-function NPI processes, across Retail, Operations and IS&T.Experience with Automation skills using Ansible, Chef, Jenkins, and PuppetGood to have knowledge on messaging queues like Kafka, Geneva, and StratosHands-on experience with CI/CD tools and building pipeline using Jenkins or any other toolStrong verbal and written communication skills and ability to coordinate with multiple technical and functional usersSelf-motivated with excellent time management skillsHigh attention to detail and you are good at finding edge casesExperience working on projects across languages and geos, multiple channelsFamiliar with Agile software processes and methodologies, and you can scope/prioritize issues accordinglyStrong analysis, problem solving, and troubleshooting techniquesAdditional skills a plus, not required:You have knowledge of iOS and are familiar with the features of the Messages applicationYou are familiar with basic Machine Learning and NLP conceptsYou've worked with Chatbots, Conversational AI or IVR systemsLanguage speaking / reading skills (Chinese or Spanish)ResponsibilitiesDrive issue triaging and troubleshooting by looking at application logs and application code in order to route to the correct team.Able to do root cause analysis, process improvements and implementation.Ensuring resilience through pro-active testing and preventive maintenance.Set up monitoring and alert mechanisms to address issues before they become problems.Identify automation opportunities to prevent incidents from recurring.Develop prototypes and execute experiments to help guide engineering efforts.Effectively document and communicate standards to platform users.Maintain in depth understanding of end-to-end experience and system architecture.Able to perform code fixes for problems and support activities such as incident trend analysis under minimum supervision

Back to Job Search