Job Description:
Experience required: 5+ years
Digital Operations Engineer
- A Digital Operations Engineer is responsible for maintaining and supporting cloud-based digital systems (including AWS cloud services and IoT platforms) to ensure their reliability, scalability, and security.
- 24/7 Incident Response: Provide Level-2 support to analyze and resolve AWS Cloud & IoT platform incidents around the clock, ensuring issues are addressed on-time to minimize downtime and that service levels are consistently met.
- This role involves analyzing and debugging incidents in production, executing runbooks and standard operating procedures, and continuously monitoring system performance to preempt issues.
- The engineer will collaborate with cross-functional teams (development, QA, product, etc.) to resolve incidents and improve operational workflows
Key Responsibilities
- Incident Analysis & Resolution: Troubleshoot and resolve L2 AWS and IoT incidents promptly, ensuring minimal downtime and SLA compliance.
- System Monitoring & Performance: Monitor cloud and IoT systems, respond to alerts, and take proactive actions to maintain uptime.
- SLA Maintenance & Reporting: Ensure SLA adherence by managing incidents efficiently and reporting on operational metrics.
- Runbook Execution & Documentation: Execute and enhance runbooks while maintaining clear documentation for operational processes.
- Collaboration & Escalation: Partner with cross-functional teams to resolve complex issues and manage escalations effectively.
- Continuous Improvement Initiatives: Automate and optimize operational workflows to enhance system reliability and reduce incident volume.
- Round-the-Clock Support: Provide 24/7 support through shift rotations, ensuring rapid response to critical incidents at all times.
Requirements
- Strong problem-solving skills with the ability to quickly identify root causes in complex cloud and IoT environments.
- High attention to detail and strict adherence to operational processes and runbooks.
- Clear and effective communicator with strong teamwork and collaboration abilities.
- Resilient and reliable under pressure, especially in critical incident operational scenarios.
- Eager to learn and adapt to new technologies, tools, and industry best practices.
- Service-oriented mindset focused on uptime, performance, and proactive risk mitigation.
Education & Experience
- A bachelor’s degree in computer science, Information Technology, Electronics Engineering, or a related field is required. An equivalent combination of education and experience in a technical field will also be considered.
- 5+ years of hands-on experience in IT operations, cloud infrastructure support, or a similar technical support role. Experience should include supporting production environments on AWS (or other major cloud platforms), with a track record of incident handling and system maintenance.
Technical Skills & Expectations
- Cloud Infrastructure (AWS): Understanding of AWS cloud services and architecture (compute, storage, databases, networking). Knowledge of cloud computing concepts, high availability, and security best practices.
- IoT & Edge Technologies: Familiarity with IoT concepts and platforms. Experience with connected devices or sensor data is a plus.
- Monitoring & Alerting: Hands-on experience with monitoring and logging tools (e.g., Amazon CloudWatch, Splunk). Ability to set up dashboards, alerts, and fine-tune monitoring systems (metrics, logs, traces). Respond promptly to anomalies or threshold breaches.
- AI Tool Usage: Familiarity with AI tools and prompt engineering. Hands-on experience with AWS AI services (AI Ops, DevOps Guru, Agentic AIs).
- DevOps & CI/CD: Understanding of DevOps practices and CI/CD pipelines.
- Incident Management Tools: Experience with IT service management and ticketing systems (e.g., Jira, ServiceNow) Ability to log incidents, document resolutions, and manage problem workflows. Knowledge of incident escalation and priority definitions (P1, P2).
- Documentation: Ability to create and maintain operational documentation.
- Security Awareness: Understanding of security and compliance in operations (e.g., access keys, certificates, security groups). Ability to follow security guidelines and fix vulnerabilities promptly (e.g., expired certificates, misconfigured firewalls).
- Certifications:
- AWS Certification (e.g., SysOps Administrator, Solutions Architect) preferred.
- ITIL or IoT-related certifications are a plus but not mandatory.
|
|