Tech Job Finder - Find Software, Tech Sales and Product Manager Jobs.
Sign In
OR continue with e-mail and password
E-mail address
Password
Don't have an account?
Reset password
Join Tech Job Finder
OR continue with e-mail and password
Username
E-mail address
Password
Confirm Password
How did you hear about us?
By signing up, you agree to our Terms & Conditions and Privacy Policy.

Software Engineer, Model Routing & Inference

at Cursor.ai

Back to all Python jobs
Cursor.ai logo
Industry not specified

Software Engineer, Model Routing & Inference

at Cursor.ai

Mid LevelNo visa sponsorshipPython

Posted 20 hours ago

No clicks

Compensation
Not specified

Currency: Not specified

City
New York City, San Francisco
Country
United States

**Software Engineer, Model Routing & Inference** - Join Cursor's Foundations team in New York or San Francisco to automate coding. As a Software Engineer, you'll build and enhance our AI inference platform, powering every interaction in our product. Own the entire inference path, optimizing speed, reliability, and cost-effectiveness at scale. Key responsibilities include designing and implementing the inference gateway, managing API semantics, and ensuring seamless model onboarding. Requires 3+ years of relevant experience and proficiency in Python, C++, and cloud services (AWS/GCP/Azure). Familiarity with AI, machine learning, and models is essential. Collaborate with a talent-dense, creative, and truth-seeking team to ship code and drive innovation.

Department: Engineering

Team: Foundations

Location: New York

Additional Locations: San Francisco

Employment Type: FullTime

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

About the Role

As a Software Engineer on the Model Routing & Inference team at Cursor, you'll build the inference platform that powers every AI interaction in the product.

This team owns the full inference path: making Cursor's AI faster, more reliable, and more cost-effective at a scale few teams in the world get to operate at. Every agent session, every tab completion, and every chat message flows through your stack.

Example projects include...

  • Building and evolving our inference gateway, a single abstraction over every provider's API semantics, so model onboarding becomes a config change.

  • Designing intelligent cross-provider failover so no single provider outage causes user-visible degradation.

  • Designing routing backpressure and admission control so traffic spikes don't cascade into providers.

You may be a fit if

  • You have deep experience building high-throughput, low-latency distributed systems, especially in inference serving, traffic routing, or real-time data pipelines.

  • You're comfortable reasoning about cost/performance tradeoffs at scale (GPU utilization, provider economics, capacity planning).

  • You have strong software engineering fundamentals and enjoy shipping production systems that handle millions of requests.

  • You make good calls in the gray area: weighing reliability, cost, latency, and user experience when there isn't a single "right" answer.

Applying

If there appears to be a fit, we'll reach to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.

#LI-DNI

Software Engineer, Model Routing & Inference

at Cursor.ai

Back to all Python jobs
Cursor.ai logo
Industry not specified

Software Engineer, Model Routing & Inference

at Cursor.ai

Mid LevelNo visa sponsorshipPython

Posted 20 hours ago

No clicks

Compensation
Not specified

Currency: Not specified

City
New York City, San Francisco
Country
United States

**Software Engineer, Model Routing & Inference** - Join Cursor's Foundations team in New York or San Francisco to automate coding. As a Software Engineer, you'll build and enhance our AI inference platform, powering every interaction in our product. Own the entire inference path, optimizing speed, reliability, and cost-effectiveness at scale. Key responsibilities include designing and implementing the inference gateway, managing API semantics, and ensuring seamless model onboarding. Requires 3+ years of relevant experience and proficiency in Python, C++, and cloud services (AWS/GCP/Azure). Familiarity with AI, machine learning, and models is essential. Collaborate with a talent-dense, creative, and truth-seeking team to ship code and drive innovation.

Department: Engineering

Team: Foundations

Location: New York

Additional Locations: San Francisco

Employment Type: FullTime

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

About the Role

As a Software Engineer on the Model Routing & Inference team at Cursor, you'll build the inference platform that powers every AI interaction in the product.

This team owns the full inference path: making Cursor's AI faster, more reliable, and more cost-effective at a scale few teams in the world get to operate at. Every agent session, every tab completion, and every chat message flows through your stack.

Example projects include...

  • Building and evolving our inference gateway, a single abstraction over every provider's API semantics, so model onboarding becomes a config change.

  • Designing intelligent cross-provider failover so no single provider outage causes user-visible degradation.

  • Designing routing backpressure and admission control so traffic spikes don't cascade into providers.

You may be a fit if

  • You have deep experience building high-throughput, low-latency distributed systems, especially in inference serving, traffic routing, or real-time data pipelines.

  • You're comfortable reasoning about cost/performance tradeoffs at scale (GPU utilization, provider economics, capacity planning).

  • You have strong software engineering fundamentals and enjoy shipping production systems that handle millions of requests.

  • You make good calls in the gray area: weighing reliability, cost, latency, and user experience when there isn't a single "right" answer.

Applying

If there appears to be a fit, we'll reach to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.

#LI-DNI

SIMILAR OPPORTUNITIES

No similar jobs available at the moment.