Atlas creates self-service tools that help you build, manage, and scale human annotation workflows while reducing operational costs through automation. Atlas is an ecosystem of services designed to efficiently support human-in-the-loop (HiTL) needs across all modalities. Our platform currently supports hundreds of daily active AI Trainers processing millions of labels per month across critical programs including Rufus, Buy for Me, and Search.We're seeking a Software Development Engineer II to help build and scale Atlas. You'll work on composable building blocks that customers can combine into complex workflows, developing features that directly impact how AI systems are trained and improved across Amazon. This role offers the opportunity to work on challenging problems at scale, combining full-stack development with distributed systems architecture.Key job responsibilities* Design, develop, and deploy full-stack features for Atlas using modern web technologies and AWS serverless architecture* Build scalable backend services using AWS Lambda, Step Functions, EventBridge, DynamoDB, and SNS/SQS* Develop responsive, high-performance user interfaces using TypeScript, React, and Redux* Build self-service capabilities through intuitive user interfaces for SOP authoring, workforce management, and workflow composition* Develop highly customizable in-line telemetry directly within the Atlas UI* Collaborate with cross-functional teams to onboard new programs and reduce operational expenditure through automation* Participate in code reviews, design discussions, and operational excellence initiatives* Own features end-to-end, from design through deployment and monitoring
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.
One Confluent. One Team. One Data Streaming Platform.
About Product & Team:
Kafka Connect is the foundational streaming integration framework for Apache Kafka. It enables reliable, scalable data movement between Kafka and the broader ecosystem of databases, cloud services, and storage systems. Designed for mission-critical production environments, this highly configurable system provides a distributed and fault-tolerant architecture that supports high-volume data pipelines across hybrid and cloud-native deployments. As a strategic core of the Confluent Cloud platform, our team builds the extensible engine that transforms and routes data streams at scale, powering the infrastructure that our customers rely on for their most critical data workflows.
What You’ll Do:
- Design, build, and operationalize high-performance, scalable, and resilient distributed systems for the Kafka Connect platform.
- Independently own the end-to-end execution of complex technical projects, ensuring high-quality delivery in production.
- Solve complex challenges in the stream processing space using cloud and on-premises resources to improve the performance, scalability, and elasticity of Kafka Connectors.
- Engage with the Open Source community to provide technical guidance and thought leadership.
- Promote a culture of engineering excellence through clear communication, high-quality execution, and effective collaboration.
What We're Looking For:
- 6 to 10 years of experience in backend engineering.
- Strong fundamentals in distributed systems design and development
- Experience in building and operating large-scale production services in public cloud environments (AWS, Azure, or GCP).
- Solid understanding of relational or non-relational database systems and systems operations (networking, operating systems, etc.).
- Proficiency in Java, Go, or C/C++, and experience with Kubernetes and Terraform.
- Understanding of systems operations (databases, networking, operating systems).
- A self-starter with the ability to work effectively in a team and excellent written and verbal communication skills.
- Bachelor’s degree in Computer Science or a related field, or equivalent experience.
Who you are:
- Empathetic and collaborative team player.
- Strong sense of ownership and commitment to team and company success.
- Self-starter driven by the challenges of a fast-paced software environment.
What gives you an edge:
- Experience in open-source development.
- Experience using Apache Kafka.
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.
One Confluent. One Team. One Data Streaming Platform.
What You Will Do:
- Build mission-critical infrastructure: Design, develop, and operationalize high-performance, scalable, and resilient backend services that form the core of the Connect Cloud Platform.
- Develop the execution engine: Build the orchestration, auto-scaling, and lifecycle management systems required to reliably run thousands of distributed Kafka Connect workloads in a multi-tenant cloud environment.
- Ensure robust isolation & security: Solve complex distributed systems challenges to ensure strict resource isolation, security, and performance guarantees for diverse workloads running on our shared platform.
- Troubleshoot & debug: Dive deep into a complex technical stack that includes microservices, Kubernetes, containers, and virtualization to resolve critical issues.
- Drive operational excellence: Ensure the operational readiness of our services and consistently meet availability and performance SLA commitments to our customers.
- Innovate & improve: Act as a strategic engineer who constantly identifies and implements architectural, process, and operational improvements.
What You Will Bring:
- 8+ years industry experience designing, building, scaling and supporting backend systems in production with a solid grasp on good software engineering practices such as code reviews, deep focus on quality, and documentation.
- Strong programming and algorithmic skills. Proficiency in a major programming language, e.g. Java, Go, C / C++, Python, etc.
- Experience configuring and deploying distributed systems and microservices using modern tools, e.g. Kubernetes, Helm, etc.
- Strong focus on project delivery and communication skills.
- Experience in driving operational excellence for large, production services.
- A strong sense of customer centricity, teamwork, technical leadership and mentorship and are excited about team and company success.
- BS, MS, or PhD in computer science or a related field, or equivalent work experience.
What Gives You an Edge:
- Proven track record of delivering large-scale, highly available, low latency, high quality systems.
- Hands-on technical expertise in large scale systems engineering or distributed systems.
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.
One Confluent. One Team. One Data Streaming Platform.
About the Role:
Confluent Cloud processes millions of events per second across AWS, GCP, and Azure. When incidents happen in a multi-cloud streaming platform, they happen at scale—data in motion, exactly-once semantics, and cascading failure modes that require deep systems thinking. We need an expert-level engineer who can drive proactive reliability improvements that prevent these incidents before they occur.
This role combines hands-on technical work with strategic program ownership. You'll spend roughly 75% of your time on engineering: building automation, improving tooling, analyzing systemic failure patterns, and designing reliability improvements. The remaining 25% is teaching and coordination: coaching teams through post-mortems, training incident commanders, and evolving our incident response practices.
You'll be part of a global team with follow-the-sun coverage, with clean handoffs that keep everyone working sustainable hours. Confluent has 800-1000 engineers across highly autonomous teams. This role sits within Cloud Architecture and Reliability - Supportability (CAR-S), a horizontal team that owns reliability standards and tooling across engineering. You're the person who makes us need incident management less.
What You Will Do:
Proactive Reliability Engineering (~75% of role) · Analyze systemic failure patterns and design improvements that prevent incident recurrence · Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments · Build tooling and automation to reduce incident response toil and scale team impact · Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack · Analyze reliability data to identify systemic improvements; build dashboards that drive action · Explore AI-assisted approaches to documentation quality and incident analysis · Design scalable reliability standards that reduce reactive workload over time.
Incident Management Program (~25% of role) · Own standards, practices, and continuous improvement of incident response · Serve as an on-call Incident Commander for production incidents, including acting as escalation IC when incidents exceed a team's management chain · Develop and deliver training programs for engineering teams at all levels · Coach teams through post-mortems and on developing actionable corrective actions
Customer Root Cause Analysis (CRCA) · Edit and review customer-facing incident documents to ensure quality and clarity · Drive turnaround SLAs while maintaining technical accuracy · Ensure clear explanation of what happened, why, and how we'll prevent recurrence
Cross-Team Leadership · Partner with engineering leaders to elevate reliability practices · Be the expert who teams proactively engage for guidance
What You Will Bring:
10+ years in SRE, incident management, or reliability engineering · Cloud experience with at least one of AWS, GCP, or Azure·
Deep expertise with incident management tooling (Rootly, PagerDuty, or similar platforms)
Strong understanding of distributed systems and failure modes at scale—Kafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systems
Deep experience with observability: metrics, logging, tracing—ability to diagnose complex issues · Kubernetes and container orchestration experience · Understanding of CI/CD pipelines and release processes · Systems thinking: understanding how infrastructure design choices affect failure modes and recovery · Familiarity with SLO/SLA frameworks.
Track record as a trusted advisor across engineering organizations · Experience driving org-wide process and cultural changes · Strong written communication (design docs, one-pagers, runbooks) · Post-mortem facilitation experience · Experience with async collaboration across time zones
Large company experience navigating reliability/incident programs at 500+ engineer organizations·
What Gives You an Edge:
Multi-cloud experience (minimum 2+ of AWS/GCP/Azure).
Modern CI/CD, GitHub, AI-assisted workflows—you'll have the freedom to build what you need
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.
One Confluent. One Team. One Data Streaming Platform.
About the Role:
As a Staff Application Security Engineer at Confluent, you will join a team of security architects and engineers responsible for shaping and advancing the application security strategy across our on-premises products and cloud services. In this role, you will go beyond implementation to define the long-term security posture of our ecosystem, spanning high-scale distributed systems, on-prem deployments, and globally operated cloud platforms.
You will lead the design and evolution of application security architecture, ensuring security is embedded throughout the product lifecycle—from early design decisions to cloud deployment and ongoing operations. Acting as a strategic partner to Engineering and Product leadership, you will influence architectural direction and proactively mitigate systemic and emerging security risks.
This role plays a key part in building and sustaining a strong security culture across Engineering, Product, and the broader organization. You will architect and oversee security automation and tooling that scales security operations and enables consistent, high-quality outcomes. The ideal candidate brings deep technical expertise and sound security judgment, with a proven ability to eliminate entire classes of vulnerabilities through architecture, automation, and cross-functional leadership.
What You Will Do:
- Partner closely with Engineering, Product, and Platform teams to identify security risks early, influence architectural decisions, and drive adoption of secure-by-design practices across the organization.
- Define and standardize threat modeling frameworks and security design standards, and lead security design reviews for complex, distributed systems, providing actionable architectural guidance to engineers and product managers.
- Serve as the subject matter expert (SME) for product security implementation reviews, overseeing security code reviews and API security testing while providing definitive remediation guidance.
- Architect and drive the roadmap for security automation, building scalable software security tooling to transform product security operations and vulnerability management practices.
- Design and lead the deployment of automation and orchestration frameworks that integrate security seamlessly into the cloud-native deployment pipeline.
- Proactively identify new vulnerability classes, lead research initiatives and orchestrate complex table-top exercises to keep the organization ahead of the evolving threat landscape.
- Strategically identify and deploy advanced technology controls to maximize observability and harden key attack surfaces across the ecosystem.
What You Will Bring:
- 10–12 years of hands-on Application Security experience, who can drive measurable security improvements across large-scale, distributed systems and global engineering organizations.
- Comprehensive knowledge of security fundamentals as applied to modern web applications and cloud-native platforms including secure software design and architecture, secure coding practices, common vulnerability classes.
- Ability to partner as a trusted peer with Engineering and Product leadership to embed security into the core architecture of the organization.
- Ability to lead technical investigation and response to application security incidents while driving preventive improvements through architecture and automation.
- Proven experience evolving the software development lifecycle to embed security by default, from securing CI/CD pipelines and build systems to implementing automated security guardrails in cloud-native deployment workflows. Passionate about applying AI and LLMs to automate complex security workflows, reduce manual toil, and drive measurable improvements in security outcomes.
- Experience in Go, Python, or Java, with the ability to design and build scalable security automation frameworks.
- Experience in leading cross-functional initiatives in distributed environments, translating security requirements into clear, executable technical roadmaps.
- A data-driven decision-maker who can balance security requirements with business velocity and engineering trade-offs to deliver outcomes.
- Ability to raise the organization’s security bar through architectural reviews, advanced technical guidance, and the development of engineers across all levels.