Jobgether

Staff Engineer - Distributed Systems

Jobgether

Node.jsTypeScriptGoGKEGCP Pub/SubCloud TasksRedisMongoDBFirestoreElasticsearch

About the Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer - Distributed Systems based in India. This role focuses on the architecture and resilience of large-scale distributed systems operating under significant production load. You will look across services, queues, databases, infrastructure, and deployment patterns to identify systemic risks before they become incidents. The role begins with a high-throughput automation platform and expands across communication and emerging product systems. You will influence critical-path architecture while remaining hands-on, spending meaningful time prototyping solutions, resolving complex failures, and shipping remediation. You will help establish capacity models, consistency guarantees, isolation boundaries, and failure-handling strategies that can withstand continued growth. This is a high-ownership environment where technical rigor, proactive problem solving, and clear cross-team communication are central to the role. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer - Distributed Systems based in India. This role focuses on the architecture and resilience of large-scale distributed systems operating under significant production load. You will look across services, queues, databases, infrastructure, and deployment patterns to identify systemic risks before they become incidents. The role begins with a high-throughput automation platform and expands across communication and emerging product systems. You will influence critical-path architecture while remaining hands-on, spending meaningful time prototyping solutions, resolving complex failures, and shipping remediation. You will help establish capacity models, consistency guarantees, isolation boundaries, and failure-handling strategies that can withstand continued growth. This is a high-ownership environment where technical rigor, proactive problem solving, and clear cross-team communication are central to the role. Accountabilities: Own the architecture health of large-scale distributed systems, including failure modes, capacity constraints, consistency guarantees, resilience, and interactions across numerous services and deployments. Review and influence critical-path technical designs, providing rigorous architectural guidance and making well-reasoned recommendations when teams face complex technical trade-offs. Proactively identify systemic risks such as single points of failure, unbounded queues, missing idempotency, thundering-herd effects, capacity constraints, and potential data-loss scenarios, then drive remediation before incidents occur. Build and ship solutions for complex architectural problems, including prototypes, reliability improvements, production remediation, and fixes for difficult cross-team issues. Design resilience into systems through degradation strategies, backpressure mechanisms, isolation boundaries, capacity models, and failure-handling patterns capable of supporting sustained growth. Investigate and resolve the most challenging distributed-system failures by understanding interactions across application services, messaging, databases, infrastructure, and deployment environments. Work hands-on with Node.js/TypeScript and/or Go to prototype architectural solutions and implement critical fixes when required. Establish and improve engineering practices through design reviews, post-mortems, architectural patterns, documentation, and technical guidance that can be adopted across multiple teams. Help define how AI-assisted engineering can be used safely and effectively when developing and operating highly critical distributed systems. Work with technologies including GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch across large-scale production environments. Help engineering teams prepare systems for substantial traffic growth by identifying capacity limits, improving observability, and validating critical failure modes. Influence engineers across teams through technical rigor, collaboration, mentorship, and clear communication rather than relying solely on organizational authority. Requirements: 10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems; experience operating hundreds of services, thousands of instances, or systems processing billions of events is strongly preferred. Demonstrated experience carrying direct accountability for production systems through significant incidents, migrations, reliability challenges, and operational failures. Deep expertise in queueing and asynchronous architectures, including delivery semantics, ordering, backpressure, idempotency, and the practical limitations of exactly-once processing. Strong knowledge of multiple storage technologies across SQL and NoSQL environments, with the ability to reason about consistency models, indexing at scale, performance, and appropriate technology selection. Expert-level experience with Redis or comparable in-memory systems, including behavior under memory pressure, network partitions, failover scenarios, and other failure conditions. Strong production experience operating Kubernetes at scale, including resource limits, autoscaling, capacity planning, node failures, and workload behavior under infrastructure disruption. Fluent in Node.js and/or Go, with sufficient hands-on ability to prototype technical proposals and implement production fixes on critical paths. Exceptional technical communication skills, including the ability to produce design documents, architectural diagrams, and root-cause analyses that drive decisions across multiple engineering teams. Strong systems-thinking ability, with an instinct for evaluating tail behavior, failure modes, network partitions, capacity constraints, and high-load scenarios. Proven ability to influence engineering teams through technical depth, constructive disagreement, clear reasoning, and respect. Strong ownership, curiosity, judgment, and problem-solving skills, particularly when dealing with ambiguous or cross-functional technical challenges. Experience using AI-assisted development tools effectively on complex systems is an advantage, particularly when balancing development speed with code quality, reliability, and operational safety. GCP-native experience with technologies such as Pub/Sub, Cloud Tasks, GKE, and Firestore is a strong advantage. Previous experience in a Staff, Principal, Architect, or comparable systems-focused engineering capacity is beneficial. Experience taking new systems from initial architecture through production while also hardening mature systems is a plus. Benefits: Full-time remote opportunity for engineers based in India. Opportunity to work on distributed systems operating at substantial scale, including billions of automation actions and messages and tens of thousands of requests per second at peak. Broad technical scope spanning application runtimes, messaging, asynchronous processing, databases, caching, Kubernetes, cloud infrastructure, and observability. Opportunity to influence architecture across multiple engineering teams and critical production systems. Significant hands-on ownership, with the ability to prototype, build, ship, and remediate solutions rather than working solely in an advisory architecture capacity. Exposure to large-scale technologies including Node.js, TypeScript, Go, GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch. Opportunity to develop resilience, capacity planning, consistency, and failure-management strategies for systems experiencing continued growth. Collaborative environment where technical rigor, clear communication, proactive problem solving, and engineering mentorship are valued. Opportunity to establish engineering patterns and practices that influence a broad engineering organization. Opportunity to work with AI-assisted engineering tools and help define safe, high-quality practices for their use in critical production systems. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

You'll be redirected to Jobgether's application page

Job Details

Salary

Not disclosed

Location

India

Job type

Full-time

Category

Cloud Engineering

Experience

10+ years

Posted

Today

Job Highlights

  • 10+ years level role
  • 100% Remote — open to candidates in India
  • Full-time position

About Jobgether

This job is hosted by Jobgether. Clicking Apply opens their site.

More jobs from Jobgether on RC9

Remote Work Style

Mixed

Mix of flexible and scheduled meetings

Your Match

See how well your skills line up with this role, and what you're missing.

AI Cover Letter

Generate a cover letter tailored to this job from your profile.