Skip to main content

Agent Trajectory Evaluation: LLM-as-Judge vs. Jev-as-Judge

By Spandana Vanamala · · 5 min read
Istock 1497108072 (1)

Evaluating AI Agent trajectories with two complementary approaches

Why Trajectory Testing Matters

When an AI agent handles a task, it doesn’t just produce one output. It takes a series of actions: calling tools, reading results, and deciding what to do next. This sequence is called the trajectory.

Testing only the final answer misses a lot. An agent can arrive at the right answer through a risky or inefficient path, skipping a verification step, calling the same tool twice, or misreporting a tool result. Trajectory testing asks the more useful question: “Did the agent behave correctly along the way?”

Trajectories can be evaluated in several ways: deterministic trajectory matching, LLM-as-judge, and newer approaches like Jev-as-judge, among others. This post focuses on two of them, LLM-as-judge and Jev-as-judge, to show how differently they can score the exact same trajectory, with real code examples and real comparison results, using a simple trajectory as the running example.

The Trajectory Under Test

To illustrate both evaluation approaches, we use a simple agent built around a single tool, get_weather(city).

It’s evaluated against a small validation dataset of real user questions:

[

  "What's the weather in Seattle today?",

  "What will the weather be like in Austin, Texas, this weekend?"

]

These two questions are chosen to expose a real gap. The agent’s tool only ever returns current conditions. “Today” happens to line up with “current,” so the Seattle trajectory looks clean. “This weekend” does not, so the Austin trajectory contains a genuine bug for the evaluators to catch.

High-level code

def run_agent(user_message: str) -> list:

    trajectory = [{"role": "user", "content": user_message}]

    city = extract_city(user_message)

    trajectory.append({"role": "assistant",

        "tool_calls": [{"name": "get_weather", "args": {"city": city}}]})

    result = get_weather(city)

    trajectory.append({"role": "tool", "name": "get_weather",

        "content": json.dumps(result)})

    trajectory.append({"role": "assistant",

        "content": f"It's {result['temp_f']}°F and {result['condition']}."})

    return trajectory

Every step gets recorded into a trajectory, a plain list of role-tagged messages. That list is the only contract both judges need.

Approach 1: LLM-as-Judge

Given the full trajectory to a second LLM with a plain-English rubric and let it reason in free text about whether the agent behaved well.

High-level code

judge = create_trajectory_llm_as_judge(

    prompt=WEATHER_AGENT_RUBRIC,

    model="openai:gpt-4o-mini",

)

result = judge(outputs=trajectory)

# {"score": true, "reasoning": "Called get_weather with the

#  correct city; reply matches the tool's result exactly."}

Strength: handles nuance and gives a human-readable explanation, great for open-ended judgment calls.

Weakness: slower, costlier, and, as we found below, inconsistent from model to model on the exact same trajectory.

Approach 2: Jev-as-Judge

Given, Jev, released by TypeSafe AI in September 2026, isn’t a text generator. It takes a state plus a bounded question and returns a typed answer with a confidence score.

High-level code

result = jev.ask(

    state={"trajectory": trajectory},

    question="Does the reply's temperature match the tool result?",

    choices=["yes", "no"],

)

# {"answer": "yes", "confidence": 0.98, "latency_ms": 410}

Strength: fast (~0.44s avg), cheap (~$0.00035/call), and far more consistent run to run than an LLM judge, per LangChain’s September 2026 benchmark.

Weakness: needs the question framed as fixed choices up front, with no free form explanation.

Validating Multiple Questions on One Tool Call

A single tool call can go wrong in more than one independent way: wrong argument, malformed argument, or a correct call whose result gets misreported to the user. So instead of one vague question, we run a checklist against just that one tool call:

  1. Does the tool call’s argument match what the user asked for?
  2. Is the argument well-formed (non-empty, real city)?
  3. Does the final reply temperature match the tool result?
  4. Does the final reply condition match the tool result?
  5. Does the final reply address the time frame the user asked about, not just current conditions?

High-level code

context = extract_tool_call_context(trajectory, "get_weather")

for q in questions:

    result = backend.ask_question(context, q.question, q.choices)

    record_pass_fail(result, expected=q.expected)

Here’s the actual output from running this checklist through Jev on both trajectories in the dataset:

Seattle (“today”) passes every check. Austin (“this weekend”) fails exactly one, question 5, because the reply only reports current conditions. That’s the checklist doing its job: catching one specific, real gap instead of a vague overall “looks fine.”

Comparing Judges: Same Trajectory, Different Verdicts

Now send the Austin trajectory, the one with a genuine bug, and the same 5-question checklist to three LLM models and Jev, then compare.

Results at a glance

Question gpt-4o-mini gpt-4o claude-sonnet-4-6 jev
City argument matches user request PASS PASS PASS PASS (0.96)
City argument is well-formed PASS FAIL PASS PASS (0.90)
Reply temperature matches tool result PASS PASS PASS PASS (0.92)
Reply condition matches tool result PASS PASS PASS PASS (0.97)
Reply addresses requested time frame PASS FAIL FAIL FAIL (0.96)

The takeaway: gpt-4o-mini scored the trajectory 100%, it missed the time-frame bug entirely. gpt-4o caught the real bug but also flagged a false positive, landing at 60%. claude-sonnet-4-6 caught the real bug cleanly at 80%. Jev also caught it cleanly at 80%, with a 0.96 confidence score attached to that specific finding. Three LLM judges disagreed with each other on the exact same trajectory and exact same question, and the weakest model missed a genuine bug completely.

Head-to-Head Comparison

Agent1

 

LLM-as-Judge Jev-as-Judge
Output format Free text + score Typed answer + probability/confidence
Speed Seconds per call ~0.44s average
Cost Expensive at scale ~$0.00035/call
Consistency Can vary run-to-run / model-to-model Much lower variance, highly repeatable
Best for Nuanced judgment + explanation High-volume, bounded checks (CI)

Conclusion

LLM-as-judge and Jev-as-judge answer the same underlying question, did the agent’s trajectory do the right thing? in two different formats. LLM-as-judge reasons in free text against a rubric and explains itself; Jev-as-judge answers bounded questions with a typed response and a confidence score. On the same trajectory and the same checklist, the LLM judges we tested disagreed with each other, while Jev stayed consistent across every question.

  • Use LLM-as-judge when you need nuanced, explainable judgment on open-ended behavior
  • Use Jev-as-judge when you need fast, cheap, repeatable answers to specific, bounded questions

References