View All Jobs 135004

Lead Cloud Engineer - Remote Eligible

Lead the design and deployment of scalable, secure AWS infrastructure and DevOps practices
Remote
Senior
4 days ago
Checkmate

Checkmate

Provides middleware that integrates restaurant POS systems with online ordering and delivery platforms to sync menus, orders, and reporting.

Lead Cloud Engineer

We are looking for a Lead Cloud Engineer to serve as the technical leader of our DevOps function, with deep expertise in AWS infrastructure, reliability engineering, and operational excellence. This role is ideal for a senior-level engineer who has experience leading other team members designing and operating production systems at scale and is comfortable owning cloud architecture decisions end-to-end.

As a Lead Cloud Engineer, you will define AWS standards, guide platform architecture, and lead initiatives that improve scalability, security, performance, and cost efficiency. You will partner closely with application engineers and other engineering team members to ensure systems are production-ready, observable, and resilient, while mentoring other engineers and raising the overall DevOps and cloud maturity of the organization.

Essential Job Functions:

AWS Cloud Architecture & Platform Leadership

  • Lead the design and evolution of AWS infrastructure supporting highly available, scalable production systems.
  • Define architectural standards and best practices across AWS services such as VPC, EC2, ECS/EKS, RDS, S3, ALB/NLB, IAM, and CloudFront.
  • Lead cloud-level decision making, including trade-offs around scalability, reliability, cost, and operational complexity.
  • Drive cloud modernization initiatives and guide teams toward resilient, well-architected AWS solutions that scale.

Infrastructure as Code & Automation

  • Own and maintain infrastructure as code using Terraform and/or CloudFormation.
  • Design reusable, modular infrastructure components that enable consistency across environments.
  • Build automation for provisioning, configuration, scaling, and lifecycle management of AWS resources.
  • Eliminate manual operational tasks through scripting, tooling, and platform improvements.

Reliability, Monitoring & Incident Response

  • Define and own monitoring, logging, and alerting standards across AWS and application services.
  • Build observability using tools such as Datadog, CloudWatch, Prometheus, or equivalent.
  • Lead infrastructure-related incident response, including coordination, mitigation, and communication in achievement of RTO and RPO objectives.
  • Conduct thorough root-cause analysis and drive long-term reliability improvements.

Database Performance Optimization & Scaling

  • Partner with application engineers to ensure databases are performant, scalable, and cost-effective.
  • Optimize AWS database services such as RDS and Aurora for performance, availability, and growth.
  • Analyze and improve query performance, indexing strategies, connection management, and resource utilization.
  • Design and implement scaling strategies including read replicas, storage scaling, and high-availability configurations.
  • Monitor database performance and proactively address bottlenecks before they impact customers.
  • Support backup, recovery, and disaster-recovery strategies aligned with business requirements.

Security, Compliance & Cost Management

  • Implement AWS security best practices including IAM, network segmentation, encryption, and audit logging.
  • Support SOC 2 and other compliance efforts through secure infrastructure design and change controls.
  • Monitor and optimize AWS costs, driving efficient resource usage without compromising reliability.
  • Ensure infrastructure changes follow defined approval, review, and documentation processes.

Leadership & Cross-Functional Collaboration

  • Act as a technical leader and mentor for DevOps and platform engineers.
  • Set expectations and standards for operational excellence across engineering teams.
  • Communicate architecture decisions, system risks, and operational status clearly to technical and non-technical stakeholders.
  • Take full ownership of cloud initiatives from design through implementation and ongoing operations.
+ Show Original Job Post
























Lead Cloud Engineer - Remote Eligible
Remote
Engineering
About Checkmate
Provides middleware that integrates restaurant POS systems with online ordering and delivery platforms to sync menus, orders, and reporting.