Skip to main content
PrismCV
JobsExtensionPricing
LoginCheck Your Resume
Check Your Resume
← Back to all jobs

AI/ML Engineer

Lemma · San Francisco, CA, US

$120k - $200k
On-site
Full-time
Mid
Check your resume against this jobApply on Ycwaas

Job Description

Lemma is production monitoring for AI agents. We catch the silent failures your observability tools and evals miss (think bad tool calls, lost context, and infinite loops) before your users find them.

Why this role exists

Agents break silently. They call the wrong tool, forget what the user said three turns ago, and loop until someone pulls the plug. Soon they'll be responsible for the majority of the world’s economic work, and most teams won't even know when they fail.

Making agents reliable is the problem Lemma exists to solve. That means catching the unknown unknowns, the failures nobody thought to write an eval for, and closing the loop end to end so they get fixed, not just flagged. It's the foundation for building agents people can actually trust and the future of self-improving systems.

The hardest part of our product is deciding what counts as a failure.

There is no ground truth here, and no benchmark to climb. Every customer's agent is different, what "wrong" means changes from one to the next, and we have to get it right across production without anyone telling us what to look for. Being confidently wrong often costs us more trust than being right fifty times earns.

This role owns the intelligence in the loop: what we flag, how sure we are, and whether the fix we propose actually fixes it.

What you’ll do

  • Own detection quality. Find the failures that don't look like failures: compliant but wrong, omissions, and patterns that only show up across thousands of traces
  • Turn implicit signals into evidence. Rephrasing, abandonment, retries, and the other ways users tell you something broke without saying so
  • Build the evals for our own system. If we can't measure precision on a problem with no labels, we can't improve it
  • Make patch generation trustworthy. Reproduce the failure, verify the fix, and know when to stay quiet instead of opening a bad PR
  • Keep it affordable. LLM-as-judge on every event is easy. Doing it at a cost per event that doesn't eat the business is the actual job
  • Read real customer traces every week. The best ideas here come from staring at production, not papers

What we’re looking for

  • High slope over years of experience. New grads and dropouts welcome
  • A track record of shipping things people actually use
  • Hands-on with LLMs in production: evals, LLM-as-judge, embeddings, and knowing when a smaller model or no model is the right call
  • Real research taste. You can tell a real improvement from noise, even when there's no ground truth to check against
  • Bonus: you were the customer once. You ran agents in production and got burned

Who you’ll work with

You'll join a team of dropout founders and engineers from Amazon, Together, and Zoom. We've been founding operators at unicorns and at startups that went on to be acquired.

Onsite in San Francisco. We don’t sponsor visas.

See how well your resume matches this job before you apply

Run a free ATS check