Site Reliability Engineer, Global Banking & Markets, Vice President

Goldman Sachs
Goldman Sachs

Software Engineering

USD 150k-250k / year

Posted on Aug 5, 2026

What We Do

At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets.

Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience, and performance of core business services that underpin a global 24×7 trading operation. Working across Global Markets' front, middle, and back office functions, you will engineer reliability while balancing stringent non-functional demands for availability, latency, and resilience as well as complex, evolving business requirements.

Want to push the limit of digital possibilities? Start here.

Who We Look For

Goldman Sachs Engineers are at the forefront of innovation, driving solutions as creative collaborators in a fast-paced global environment. We seek individuals who evolve, adapt, and thrive on challenging problems.

As part of our SRE team, you will operate at the intersection of reliability engineering, cloud infrastructure, and AI-driven operations. Using Goldman Sachs' AI tooling and agentic assistants, you will accelerate incident diagnosis, automate operational toil, comprehend large legacy codebases, and raise the bar for production-quality automation across the software and reliability lifecycle. Above all, you will bring strong risk acumen and the ability to connect the right people across the organization to resolve problems quickly and decisively.

Your Impact

  • Own reliability outcomes: Define and defend Service Level Objectives (SLOs), error budgets, and reliability standards for critical trading services, with risk always front of mind.
  • Reduce risk and toil: Identify systemic risks before they materialize, automate away repetitive operational work, and strengthen the resilience posture of the platform.
  • Connect and communicate: Act as a trusted coordinator during incidents — rapidly mobilizing the right engineers, domain experts, and stakeholders across a globally distributed organization, and communicating clearly with both technical and non-technical audiences.
  • Multiply your output with AI: Orchestrate AI coding and operations agents to accelerate root-cause analysis, remediation, and automation while maintaining mastery, quality, and production fitness over all AI-generated work.
  • Build for the future: Design and operate high-availability, multi-region, event-driven services on a modern cloud-native platform, setting the reliability and architectural standard for years to come.

What You Will Do

  • Design, build, and operate high-availability, multi-region, cloud-native services with security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer.
  • Establish and manage SLIs, SLOs, and error budgets; drive blameless post-incident reviews and translate findings into durable engineering improvements.
  • Lead incident response for latency-sensitive, high-throughput trade lifecycle systems — quickly diagnosing issues, coordinating cross-functional responders, and communicating status to stakeholders.
  • Develop event-driven architectures, multi-stage processing pipelines, and optimized data paths for high-throughput trade lifecycle management.
  • Apply strong risk acumen to change management, capacity planning, and resilience testing (chaos engineering, failover, and BCP drills).
  • Partner with engineers, domain experts, and global stakeholders to understand production processes, challenge entrenched assumptions in a cloud-centric, AI-driven world, and drive modernization.
  • Multiply your impact with a modern, AI-centric toolchain, orchestrating AI agents across the SDLC and operations to rapidly comprehend large codebases, generate production-quality automation, and accelerate delivery.

Basic Qualifications

  • 8+ years of professional software / reliability engineering experience, with strong command of at least one major language (Java 17+ preferred), including concurrency, collections, and modern language features.
  • Demonstrated risk acumen — the ability to identify, quantify, and mitigate operational and technical risk in a regulated financial services environment.
  • Excellent communication and stakeholder-coordination skills — proven ability to connect the right people quickly and drive resolution across geographically distributed, technical and non-technical audiences.
  • Proven experience running high-availability production environments: SLIs/SLOs, error budgets, on-call, incident command, and post-incident reviews.
  • Strong understanding of cloud infrastructure (GCP, AWS), container orchestration (Kubernetes, Docker), and infrastructure-as-code.
  • Working knowledge of AI models and AI-assisted engineering tools (e.g., Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist), including the ability to govern AI agents, critically assess their output, and maintain quality over AI-generated work.
  • Experience building event-driven and distributed systems, including messaging platforms (e.g., Apache Kafka), delivery guarantees, and resilience strategies.
  • Strong SDLC and automation practices: version control, CI/CD pipelines, automated build/test/deploy workflows, and code quality tooling.
  • Solid observability discipline: application instrumentation, distributed tracing, structured logging, and metrics-driven operations.
  • Ability to rapidly navigate, understand, and debug large and unfamiliar codebases — with and without AI assistance.

Preferred Qualifications

Experience with a meaningful subset of the following is highly valued:

  • Reliability & Operations: Chaos engineering, capacity planning, load/performance testing, and production support in high-availability, latency-sensitive environments.
  • Frameworks & Architecture: Spring Boot, gRPC / Protocol Buffers, integration/orchestration frameworks (e.g., Apache Camel, Spring Integration), and pipeline/adapter patterns (retry, dead-letter queues, error isolation).
  • Cloud & Infrastructure: Cloud platforms (GCP, AWS), Kubernetes/Docker, JVM tuning for containerized workloads, and infrastructure-as-code (Terraform, Helm).
  • AI & Automation: Applying AI models to operational use cases — anomaly detection, log analysis, automated remediation, and agentic operations.
  • Observability & Operations: Prometheus, Grafana, OpenTelemetry, and SLO tooling.
  • Data & Performance: Data modeling, SQL/NoSQL databases, caching strategies, and performance optimization in latency-sensitive systems.
  • Security: Enterprise security patterns; authentication protocols, mutual TLS, secrets management, and certificate rotation.
  • Domain Knowledge: Equities, post-trade, or financial services experience; trade lifecycle concepts, position management, reconciliation, and multi-system migration environments.
  • Other: Asynchronous / non-blocking I/O frameworks (e.g., Vert.x, Netty), multi-region / BCP architectures, and open-source contribution experience.

Salary Range
The expected base salary for this New York, New York, United States-based position is $150,000-$250,000. In addition, you may be eligible for a discretionary bonus if you are an active employee as of fiscal year-end.

Benefits
Goldman Sachs is committed to providing our people with valuable and competitive benefits and wellness offerings, as it is a core part of providing a strong overall employee experience. A summary of these offerings, which are generally available to active, non-temporary, full-time and part-time US employees who work at least 20 hours per week, can be found here.

ABOUT GOLDMAN SACHS

At Goldman Sachs, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world.

We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has a number of opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. Learn more about our culture, benefits, and people at GS.com/careers.

We’re committed to finding reasonable accommodations for candidates with special needs or disabilities during our recruiting process. Learn more: https://www.goldmansachs.com/careers/footer/disability-statement.html

© The Goldman Sachs Group, Inc., 2026. All rights reserved.