Skip to main content Skip to footer

Before an AI answer reaches a lawyer, it has to pass thousands of checks most legal teams never see. This is how the iManage evals process decides what "good enough to trust" really means, and why a panel of practising lawyers couldn't spot the difference until they knew which answer was the machine's.

Ten lawyers sat around a shared board, ranking anonymised answers to the same set of legal questions, debating out loud what they liked about each one. Slipped in among their human-written answers, unlabelled, were responses generated by Ask iManage. For several rounds, it sat at the top of the rankings.

Then the lawyers found out which answer was the machine's. Nothing about the words changed. What happened to the ranking is the reason two of the engineers who ran that exercise, Sam Grange and Nicole Dunleavy, can talk for hours about a discipline most knowledge workers have never heard of: evals.

What is evals, actually?

Asked for a one-sentence explanation a law firm partner with zero technical background could understand, Sam Grange, principal AI engineer leading evals at iManage, doesn't hesitate: "How do you make AI systems better?"

It's a modest-sounding question for a discipline that now runs hundreds of times a week, across the test suites Sam and a team of legal engineers have built over the past two years: nearly 6,000 test cases, checking over 54,000 individual things a response has to get right, everything from whether the right documents were retrieved in the right order to whether the final answer contains what a practising lawyer would actually expect to see.

That second kind of check is the harder one, because testing AI that produces text is nothing like testing conventional software. "When you are testing code, it's usually deterministic. The behaviour will be the same every time you run it," says Nicole Dunleavy, senior software engineer at iManage. "Whereas with an LLM, it's non-deterministic, so you can constantly get different responses. To measure that, you need a system that goes beyond matching strings and instead understands meaning."

Judging the judges

Dunleavy is quick to add that this is harder than it sounds. You cannot just ask an LLM whether an answer is good. It will hedge, hallucinate, and change its mind.

So, doing something that subjective calls for what Grange calls "graders". Systems built to judge the quality of an answer, often using another LLM as the judge. That might sound circular, using AI to mark AI's homework, but the method is more rigorous than it sounds. Grange's team doesn't ask a judge model a vague "is this good?" Instead, they give each judge a detailed, multi-faceted rubric, written by subject matter experts, that spells out exactly what a good answer for that specific question, in that specific context, has to contain. It's the difference between handing an exam to a marker with no guidance and handing them a proper marking scheme.

Even a good marking scheme needs a fair marker, though, and this is where the practice gets stranger. The model doing the judging is never the model that wrote the answer. "You wouldn't let a student judge their own work," Dunleavy says. "An LLM will like its own writing style, its own layout, its own pacing, so you use a different model to make sure the quality isn't unfairly judged." Left unchecked, an AI judge tends to rate an answer more favourably when it reads like something the judge itself would have written, a quirk the team calls narcissism. So iManage draws its judges from entirely different model families than whatever is being tested, and for the answers that matter most, uses more than one judge, sometimes called a jury of judges, and checks that they agree before trusting the result. The team also deliberately tests its judges with tricky, misleading examples, answers dressed up with confident formatting or padding, to make sure a judge can be fooled by substance and not by presentation alone.

Which is, it turns out, exactly what happened to the lawyers.

Once the panel knew which answer had come from a machine, they found its tells almost immediately: the phrasing, the pacing, the slightly-too-tidy structure. It dropped straight to the bottom of the rankings. But here's the detail Grange still isn't sure what to make of. The words on the page hadn't moved. What the panel had learned wasn't that the answer was wrong. They'd learned who wrote it. And a piece of legal advice with no name behind it, no partner willing to put their reputation on it, stopped being acceptable the moment they knew that, regardless of what it said.

That's the case for evals. Evals are how iManage tries to build, deliberately, the kind of accountability a lawyer already carries by default: an audit trail back to who decided what "good" means, and proof the answer was checked against that standard before anyone saw it.

How iManage builds and runs its evals

Nothing ships without this process running first. Before a feature is even scoped, product managers and evals engineers negotiate what good enough actually means, often bringing in practising lawyers to help define it with wisdom learnt from their years of practice. Grange calls the result "expert-derived, data-driven development" with a grin.

Evals catch problems long before a customer would. While comparing search methods, the team found that a common approach, vector search, took a highly specific court case name involving a well-known ride-hailing company and quietly broadened it into a general search for similar companies, returning a wave of loosely related case law where the lawyer needed one precise citation. Testing showed that letting an LLM reason over meaning, combined with a more traditional keyword approach, kept the precision lawyers need without losing the breadth AI is good at. That result changed how the underlying search was built, before it ever shipped.

Tested without touching a single client file

There’s some concern that legal AI vendors build realism into their evaluations by drawing on real, anonymised client data. At iManage, none of it ever runs on customer data. "We never use customer data, we couldn't get access to it even if we wanted to," Grange says. Instead, the team relies on licensed and curated documents labelled by legal engineers and subject-matter experts, and, where real-world examples are unavailable, the team goes to great expense to build synthetic data checked against real documents for realism.

“Did you know if you’re not careful an LLM will on average produce a lease 50% shorter than that you see in the wild? And all the emails supporting it will quote different numbers?” Sam wryly added.

Why should a lawyer care about any of this?

Because it's the difference between a vendor you can question and one you have to take on faith.

"If someone can't talk to their evals process, and how they do data-driven iteration with expert input, I don't know what you're paying for."

That's Grange's advice for evaluating any AI vendor: be wary of headline scores. Run the actual count: the evals programme at iManage currently spans 5,690 individual test cases and more than 54,000 individual grading checks, each one an atomic, expert-written pass or fail. A serious evals programme looks like specific numbers you can ask a vendor to defend, not one score.

Evals let iManage trace every design decision back to the expert judgment that defined what "good" looks like for that task.

As Dunleavy puts it: "It's not a test-at-the-end process. We check at every step. When we identify a gap in our evaluations, we fill it before anything ships. That is our actual development work."

The gap between adoption and confidence is well documented. Recent research found that while 85 percent of organisations report piloting or implementing AI, just 17 percent say it is fully embedded into their operations, a gap the iManage Knowledge Work Benchmark Report 2026 ties to governance and trust rather than raw model capability. Trust in AI can't be declared. It has to be earned, one checkable answer at a time, and evals are the clearest record iManage has of doing that work rather than just claiming it.

Not everything is that checkable, and the team is upfront about where the line sits. Whether a document was retrieved correctly, whether a clause was extracted verbatim, whether a summary left something out, these can be tested against a clear standard, which is why the numbers above exist at all. Whether an answer is fair to the person reading it is a different kind of question, one that doesn't reduce to a rubric nearly as easily. That's simply where the method's boundary sits.

It's also the subject of the next post in this series, where we go inside how iManage tests its AI for fairness, not just accuracy. Follow the iManage blog to get it first.

Global Head of Content

Former journalist turned tech content strategist. James has spent his career finding the stories that connect — and making sure they reach the people who need to hear them.