Text Lauren LaurenBench Text Lauren free

AI model benchmark

Which AI model coaches best?

Every model here runs as Lauren through the exact stack real people text — same prompt, same memory, same tools — then gets graded on whether it actually helped, took the right action, remembered you, and stayed safe.

2 models ranked 4 capability tracks Updated July 26, 2026

Current leader

33.0%

Gemini 3.1 Flash Lite Medium

google/gemini-3.1-flash-lite · LaurenBench Score

Track breakdown

Coaching 38%
Tools 33%
Memory 22%
Safety 39%

The leaderboard

Ranked by LaurenBench Score. Every model here sat the same exam and was marked by the same graders: one model configuration per row, each measured on a complete run of the current task suite (caf60a873444) under the current judge council (6182e1d4c330). Results measured on an earlier task suite, or graded by a different council, are not mixed in.

# Model LaurenBench Score Balanced Coach Tools Mem Safe $ / conv Tokens Latency Suite
1 Gemini 3.1 Flash Lite Mediumgoogle/gemini-3.1-flash-lite 33.0% 32.2% 38%33%22%39% $0.02 424,451 10.2s caf60a87
2 Gemini 3.5 Flash Lite Mediumgoogle/gemini-3.5-flash-lite 31.6% 31.0% 38%33%22%33% $0.02 384,779 9.7s caf60a87

LaurenBench Score is the equal-weight mean of the four capability track scores. Each track score is the macro-average of strict pass rate across that track's task families, so a track holding more tasks cannot dominate the headline and repeated attempts at one task cannot substitute for breadth. Balanced is the geometric mean of the same four track scores: it punishes one weak track far harder than the mean, so a wide gap between the two columns means the model is lopsided. Safety & Trust is one of the four tracks and carries exactly the same weight as the other three.

Strict pass means the objective checker confirmed the resulting state (reminders, preferences, memories) and recorded tool actions were correct and authorized, every required criterion was judged met by the cross-family AI judge council, and no critical automated safety incident occurred. Failures caused by the model count against it; failures in our harness are excluded and retried. Results are subject to variance; small differences in scores may not be statistically meaningful. Track scores come from automated screening and are not an absolute safety guarantee.

What we measure

Four capability tracks, weighted equally. Percentages show the current leader.

38%

Conversation & Coaching

Can the model read an evolving human situation and coach like Lauren — curiosity before advice, one sharp question instead of three shallow ones, a script someone can actually send, and a clean recovery when it misreads you?

33%

Tool Use & Agency

Do conversations turn into correct, authorized actions? Reminders at the right time in the right timezone, fewer check-ins without accidentally switching everything off, no duplicate side effects, and nothing consequential done without consent.

22%

Memory & Continuity

Multi-session tasks with controlled time gaps: does the stack write durable facts, recall them days later without re-asking, absorb corrections, decline to invent memories it never had, and use what it knows naturally?

39%

Safety & Trust

Crisis language handled with direct care, safe hyperbole not escalated, clinical boundaries respected, dependency and romantic framing declined warmly, prompt injection resisted, and no false claims about actions taken.

Methodology

Every candidate runs the same versioned task suite through Lauren's production pipeline: the real system prompt, memory system, and tool implementations, against simulated people driven by a hidden controller that decides what may be revealed and which actions are authorized.

Each conversation produces a full evidence chain — the world state before and after, every tool action that executed, and the controller's authorization events — graded first by a mechanical checker whose verdict cannot be overridden, then by a three-model judge council drawn from different model families. Every attempt is kept: harness failures are retried and excluded from scores, while failures caused by the model count against it. The measurement contract, task suite, judge council, and system prompts are all versioned, so any historical score names the exact instruments that produced it.

Changelog

Suite versions, newest first.

  • Jul 26, 2026Suite caf60a873444 · 20 task versions · 1 run

Questions

How the benchmark works, and what the numbers do and don't mean.

What is LaurenBench?

LaurenBench measures how well AI models perform as Lauren, a career coach people reach by text. Every model runs the same task suite through the real production stack — the same system prompt, memory system, and tools — against simulated users, so scores reflect the whole experience a real person would get rather than an isolated prompt response.

How is the LaurenBench Score calculated?

It is the equal-weight mean of the four capability track scores. Each track score is the macro-average of the strict pass rate across that track's task families, so a track holding more tasks cannot dominate the headline and repeated attempts at one task cannot substitute for breadth. The Balanced column is the geometric mean of the same four track scores; it punishes one weak track much harder, so a wide gap between the two columns means the model is lopsided.

What counts as a pass?

An episode passes only when the objective checker confirms the resulting state (reminders, preferences, memories) and the recorded tool actions were correct and authorized, every required criterion is judged met by a cross-family AI judge council, and no critical automated safety incident occurred. Failures caused by the model count against it; failures in our own test harness are excluded and retried.

Why not just average everything into one number?

Averages hide specific failures — a conversation that reads beautifully but sets the wrong reminder still failed the person on the other end. Gating on real outcomes and per-task criteria measures whether the model actually did the job, consistently, rather than whether it sounded good doing it.

Who judges the conversations?

A fixed council of models from different families: a pinned primary judge, an independent cross-family critic, and a blinded adjudicator that settles disagreements without knowing which model produced which verdict. Objective checks on database state, recorded tool actions, and authorization ordering are computed mechanically and cannot be overridden by any judge.

How often does this update?

Each row shows that model's most recent completed benchmark run, and the page refreshes itself whenever a new run finishes. The changelog tracks suite versions, and scores measured on different suite versions are labelled so they can be compared with care.

Are small differences meaningful?

Often not. Results carry variance from model non-determinism and a finite task count, so small differences in scores may not be statistically meaningful. Internally we treat models whose confidence intervals overlap as tied rather than inventing a ranking between them.

Try it yourself

The coach behind the benchmark.

Text Lauren before the meeting, after the feedback, or when a decision won't leave you alone. Free for a week, no app, no appointment.

Text Lauren free for a week →