Judgment Labs Raises $32M to Build the Improvement Layer for AI Agents

6 min read
Co-founder of JustAINews
Share
Key Points
  • Judgment Labs raised $32 million to help AI teams monitor, evaluate, and improve agents directly from production data.
  • Lightspeed Venture Partners led both rounds, joined by Nova Global, SV Angel, Valor Equity Partners, and Dynamic, plus Stanford Professor Chris Manning and the founders of DoorDash and Mercor.
  • The platform tracks full agent trajectories to surface recurring failure patterns and turn real interactions into concrete product improvements.
Alex Shan, co-founder and CEO of Judgment Labs
Credit: Judgment Labs official website

Judgment Labs has closed $32 million in combined seed and Series A funding to build infrastructure for improving AI agents from production data. The San Francisco company was founded by three childhood friends: CEO Alex Shan, Chief Scientist Andrew Li, and CTO Joseph Camyre, all in their early twenties. Their core argument is simple: the tools most teams use to measure agent quality were not built for how agents actually behave.

The timing reflects a real shift in the market. AI agents are software systems that can reason through open-ended problems, call tools, and complete multi-step tasks with little human input. According to Grand View Research, the global AI agents market was estimated at $7.63 billion in 2025 and is projected to reach $182.97 billion by 2033, at a CAGR of 49.6%. As more companies put agents directly in front of customers, measuring whether those agents actually work has become a harder and more pressing problem.

Lightspeed Venture Partners led both rounds, returning less than six months after its initial seed investment. Nova Global, SV Angel, Valor Equity Partners, and Dynamic also participated.

Why Measuring Agent Quality Is Harder Than It Looks

To understand what Judgment Labs is building, it helps to understand what changed. For years, most AI software was built around a simple pattern: a user sends a message, the model sends one back. Measuring quality in that setup is relatively straightforward. You look at the input, you look at the output, and you judge whether the answer was good.

A newer category of AI software works differently. The company calls these "deep agents." Systems like Anthropic's Claude Code, OpenAI's Codex, and Cognition's Devin do not just respond to questions. They plan, write and run code, browse the web, ask follow-up questions, and can work on a single task for minutes or hours. The output is not one answer but a long chain of decisions, searches, and actions taken along the way.

When an agent fails, the final answer often looks only slightly wrong. The actual mistake might be buried four or five steps back: a search query phrased incorrectly, a step skipped, a guess taken when the agent should have asked for clarification. Checking the final output tells you something went wrong. It rarely tells you where.

"We set out to build Judgment because the teams building deep agents didn't have tools that understood what their agents were actually doing. Input-output evals miss so much of where agents go wrong. Lightspeed has been the right partner from day one: they backed us when we were a handful of researchers with a thesis, and they're doubling down now that the thesis is playing out in production."

Alex Shan, CEO and co-founder of Judgment Labs

How Judgment Labs Will Deploy the New Funding

Most of the capital will go toward hiring AI researchers and engineers in San Francisco. The company also plans to grow the forward-deployed engineering team that works on-site with customers.

Judgment Labs is already running in production at a number of agent-native companies, where it handles monitoring and improvement cycles for live agents. One early customer, Aqil Naeem of E3 Group, described what it replaced:

"We tried other tools, but none of them could automatically point toward where things failed. Judgment is in a different league; we can see exactly where our agents make mistakes, fix them, and measure the lift. It's the difference between guessing and knowing, and it's showing up directly in our customer experiences."

Aqil Naeem, CEO of E3 Group

The company's stated goal is to give any team building agents the tools to make those products measurably better over time, with every real interaction feeding into the next improvement.

The Founders and How the Platform Works

The three founders have known each other since childhood. Andrew introduced Alex to natural language processing when they were kids. Alex was the first paying customer of a Python course Joseph ran in middle school. Years later, each went deep into the field through a different path: Alex researched NLP at Stanford under Professor Chris Manning, Andrew was an early hire at AI training startup TogetherAI, and Joseph built large-scale infrastructure at Datadog.

They started Judgment Labs when the first wave of deep agents reached production and the failure patterns they had each studied separately started showing up in real systems. The platform works by giving teams a full view of the trajectory an agent takes across a task. It identifies patterns that repeat across thousands of interactions and converts those findings into concrete fixes teams can ship. The aim is to replace manual trace-reading with a systematic process where production data does the work.

The Investors Behind the Round

Lightspeed Venture Partners led both the seed and the Series A. Founded in 2000, the firm has spent over two decades backing companies from the earliest stages through to Series F and beyond, with a particular focus on high-conviction bets on founders with ambitious technical ideas. It manages over $40 billion in assets across its funds.

Lightspeed has a long track record in enterprise software and AI infrastructure. Its portfolio includes Anthropic, Databricks, Glean, Mistral, Rubrik, Snap, Stripe, Rippling, and Wiz, among others. The firm backed Judgment Labs at seed and chose to lead the Series A less than six months later, an unusually fast return to the same company.

"Judgment is solving the hardest problem in the agent stack — how do you measure and improve something that thinks, plans, uses tools, and remembers? The Judgment team has been productizing agentic evaluations long before the word 'evals' became popular. They have a clear technical vision, a product that agent-native startups are already standardizing on, and a market opportunity that grows every time another company puts an agent into production. We led the seed because the bet was obvious, and we led the Series A because the results have been extraordinary."

James Alcorn, Partner at Lightspeed Venture Partners

Additional participants in the round include Nova Global, SV Angel, Valor Equity Partners, and Dynamic. The company's blog announcement also notes participation from Stanford Professor Chris Manning and the founders of DoorDash and Mercor.

Funding details

Subscribe to JustAINews!

Get the industry's biggest AI news straight to your inbox.
Subscribe

Related posts

© Copyright 2025 - Just AI News - All Rights Reserved
linkedin facebook pinterest youtube rss twitter instagram facebook-blank rss-blank linkedin-blank pinterest youtube twitter instagram