Senior Software / Site Reliability Lead Engineer

General Dynamics Mission Systems · via Himalayas ·

Budget / Salary$142,696–158,303
TypeFull-time job
LocationUnited States
Posted3 hours ago
Basic Qualifications
Bachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.CLEARANCE REQUIREMENTS: Ability to obtain a Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U.S. Government security investigation and must meet eligibility requirements for access to classified information. Due to the nature of work performed within our facilities, U.S. citizenship is required.
Responsibilities for this Position
What You Will Own

Cross-pod reliability standards. Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes. You own the engineering signal; together you tell the full reliability story.

SLOs and reliability metrics. Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.

Monitoring and observability. Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do.

Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong. Alerts you build catch real problems; they don't cry wolf.

Incident response. Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog. When something breaks, you coordinate the response and ensure it doesn't break the same way again.

Production Readiness. Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship."

Toil elimination. Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.

What You Won't Own

Infrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works

Business process decisions or backlog prioritization

Business-side reliability metrics - you partner with the Functional SRE on those, but they own that domain

What Makes This Role Different

AI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.

You are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.

Your production readiness criteria directly determine whether AI services go live. You have real authority to say "not ready."

You operate across projects simultaneously — embedded deeply enough to understand large-scale systems, while maintaining consistent standards across all projects.

Your software engineering background means you can engage directly with development teams at the design level — catching reliability problems before they become operational ones.

Required Qualifications

Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience

Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines

Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.

Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)

Experience with containerized environments — Docker, Kubernetes, container orchestration at scale

Experience defining and managing SLOs, error budgets, and incident response procedures in production

U.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.

Preferred Qualifications

Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines

Software engineering fundamentals — you can read, write, and meaningfully review production-quality code. You understand how architectural and design decisions made early translate into operational problems later.

Software design experience — you have participated in or led design reviews, defined service interfaces or APIs, and pushed back on design decisions using reliability and operability as criteria

Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that have caught real problems.

Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)

Experience with containerized environments — Docker, Kubernetes, container orchestration at scale

Experience defining and managing SLOs, error budgets, and incident response procedures in production

What Sets You Apart

You build things that work. Your default response to a problem is code, not a document.

You have shipped AI systems that real users depended on in production.

You are comfortable working without detailed specs — you can take a problem statement and figure out the right approach.

You care about reliability as much as capability — you monitor what you deploy.

You move fast without being reckless. You know when to iterate and when to get it right the first time.

Details

Remote — 100% telework

9/80 schedule

Defense industry experience is not required

Salary Note
This estimate represents the typical salary range for this position based on experience and other factors (geographic location, etc.). Actual pay may vary. This job posting will remain open until the position is filled.Combined Salary Range
USD $142,696.00 - USD $158,303.00 /Yr.Company Overview
General Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!
Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans
Originally posted on Himalayas
site-reliability-engineering sre-lead software-engineer devops-engineer reliability-engineering site-reliability-engineering-lead staff-site-reliability-engineer-(sre) senior-sre-engineer senior-reliability-engineer senior-site-reliability-engineering-architect
Apply on Himalayas →

Job sourced from Himalayas. Applications happen directly on the original platform — we never collect your data.