Senior Software / Site Reliability Lead Engineer

Posted Date 6 hours ago(8/10/2026 5:25 PM)
ID
2026-74192
Job Location
US-Telework-Telework
Required Clearance
Secret, obtainable within reasonable time based on requirements
Category
Engineering-Software
Employment Type
Full Time
Hiring Company
General Dynamics Mission Systems, Inc.

Basic Qualifications

Bachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.

CLEARANCE REQUIREMENTS: Ability to obtain a Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U.S. Government security investigation and must meet eligibility requirements for access to classified information. Due to the nature of work performed within our facilities, U.S. citizenship is required.

Responsibilities for this Position

What You Will Own

  • Cross-pod reliability standards. Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes. You own the engineering signal; together you tell the full reliability story.
  • SLOs and reliability metrics. Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.
  • Monitoring and observability. Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do.
  • Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong. Alerts you build catch real problems; they don't cry wolf.
  • Incident response. Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog. When something breaks, you coordinate the response and ensure it doesn't break the same way again.
  • Production Readiness. Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship."
  • Toil elimination. Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.

What You Won't Own

  • Infrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works
  • Business process decisions or backlog prioritization
  • Business-side reliability metrics - you partner with the Functional SRE on those, but they own that domain

What Makes This Role Different

  • AI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.
  • You are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.
  • Your production readiness criteria directly determine whether AI services go live. You have real authority to say "not ready."
  • You operate across projects simultaneously — embedded deeply enough to understand large-scale systems, while maintaining consistent standards across all projects.
  • Your software engineering background means you can engage directly with development teams at the design level — catching reliability problems before they become operational ones.

Required Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience
  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production
  • U.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.

Preferred Qualifications

  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines
  • Software engineering fundamentals — you can read, write, and meaningfully review production-quality code. You understand how architectural and design decisions made early translate into operational problems later.
  • Software design experience — you have participated in or led design reviews, defined service interfaces or APIs, and pushed back on design decisions using reliability and operability as criteria
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that have caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production

What Sets You Apart

  • You build things that work. Your default response to a problem is code, not a document.
  • You have shipped AI systems that real users depended on in production.
  • You are comfortable working without detailed specs — you can take a problem statement and figure out the right approach.
  • You care about reliability as much as capability — you monitor what you deploy.
  • You move fast without being reckless. You know when to iterate and when to get it right the first time.

Details

  • Remote — 100% telework
  • 9/80 schedule
  • Defense industry experience is not required

Salary Note

This estimate represents the typical salary range for this position based on experience and other factors (geographic location, etc.). Actual pay may vary. This job posting will remain open until the position is filled.

Combined Salary Range

USD $142,696.00 - USD $158,303.00 /Yr.

Company Overview

General Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!


Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans

Apply

Sorry the Share function is not working properly at this moment. Please refresh the page and try again later.
Share on your newsfeed

Need help finding the right job?

We can recommend jobs specifically for you! Click here to get started.