About Traversal
Traversal is the AI Site Reliability Engineer (AI SRE) for the enterprise.
Production complexity was already outpacing what engineering teams could manage manually, and AI-generated code is accelerating that gap. Traversal is built for that challenge, autonomously understanding and reasoning across even the largest, most complex production environments to diagnose, fix, and prevent incidents. Our mission is to free engineers from endless firefighting and give them more time to focus on creative, high-impact work.
Today, Traversal operates in mission-critical environments at some of the world’s largest enterprises. Our roots remain deeply embedded in AI research, and we’ve brought together researchers from institutions including MIT, Harvard, Berkeley, Columbia, and Cornell with world-class technical staff and operators from companies like Google, Meta, Datadog, ServiceNow, and Citadel Securities to take on one of the hardest problems for AI to solve. Traversal is backed by Sequoia Capital, Kleiner Perkins, Hanabi, NFDG, and American Express Ventures.
The Role
As an AI Platform Engineer at Traversal, you’ll work on the core foundations that make Traversal’s AI possible while ensuring Traversal’s platform, products, and applications are delivered with a high bar for quality, reliability, resilience, cost-effectiveness, and maintainability.
You will build and own the frameworks, shared libraries, application foundations, observability, and developer tooling that power Traversal’s systems and AI agents. This involves designing scalable distributed systems and abstractions, delivering core application framework components, implementing best practices for software architecture and software delivery while balancing velocity, research flexibility, and production reliability.
Responsibilities
Lead the design and implementation of scalable, robust backend systems to support AI agents and observability tools.
Design and implement high-performance APIs to enable seamless communication between backend systems and frontend interfaces.
Architect scalable distributed systems to support real-time workloads over petabytes of heterogeneous telemetry data.
Build live evaluation pipelines, automated scoring systems, and benchmarks to measure and drive AI performance.
Collaborate with AI engineers and scientists to integrate AI-driven insights and solutions into backend systems.
Monitor, optimize, and scale backend services to handle high volumes of data while ensuring low-latency performance.
Requirements
Strong system design skills for distributed systems.
Proven production-scale software engineering experience.
Experience with LLM-based applications and/or multi-agent systems.
Strong data modeling skills and a track record of writing clean, maintainable code.
Collaborative, impact-driven mindset and ability to work across research and engineering teams.
Nice to Have
Knowledge of software incidents and production SRE workflows.
Prior experience with AI benchmarking or evaluation systems.
Experience creating quantitative scoring systems or benchmarks in new problem domains.
Familiarity with observability stacks (logs, metrics, traces) and telemetry systems.
Background in agentic architectures, orchestration frameworks, or applied AI research.
Compensation
We offer competitive compensation, startup equity, health insurance, and additional benefits. The U.S. base salary range for this full-time, in-person role in New York is $150,000–$300,000, plus equity and benefits. Our salary ranges are based on location, level, and role. Individual compensation is determined by experience, skills, and job-related knowledge.
Why You Should Join Us
Traversal is a place to take on hard, meaningful problems with real ownership from day one. You’ll work alongside people who challenge you to grow, learn constantly, and help define a new category of infrastructure software. We think long term, move quickly, and hold a high bar without taking ourselves too seriously.
We offer competitive salary and equity packages, health insurance, fertility benefits, a great tech setup stipend and flexible time off. Plus in-office snacks, team happy hours and outings, an annual company offsite, and plenty of built in time to collaborate across teams.
Traversal is fully in-office, 5 days a week, based in New York near Madison Square Park. We have a collaborative, hard-working culture and are energized by building the future of AI-powered software maintenance.