Skip to content
unlob

Trust & accuracy

Citations that do not support the answer

Citation precision — the share of citations that actually support the sentence they are attached to — is the failure users notice first and teams measure last. It shows up in review as an answer that looks well sourced until someone clicks.

Verified 21 September 2026

Three ways a citation fails

Topical match. The passage is about the right thing and asserts something else: the same company in a different quarter, the same drug for a different indication, the same regulation before it was amended. Ranked by relevance it looks ideal, which is exactly why it was cited.

Copy as confirmation. The answer cites three pages that are one wire story republished. Each citation is accurate; together they overstate the support by a factor of three, and a reader counting links is misled.

Stale support. The page supported the claim when it was published and has since been superseded. The citation is accurate about the past and wrong about the present.

None of these is fixed by telling the model to cite more carefully. The model can only cite what it was given, and in all three cases what it was given looked like support.

Check the evidence before the model writes, not the citations after

The usual remedy is a second pass: a judge model reads each citation against its sentence and flags the ones that do not hold. It works, it costs a second model call per answer, and it finds the problem after the answer was written — sometimes after the user read it.

The cheaper check sits earlier. If the retrieval layer reports what each piece of evidence is — its source role, its origin, its date — and whether the set as a whole clears a bar you set, the model writes from a set that is not structurally misleading, and the cases that would have produced a hollow citation never reach generation.

Be clear about the division of labour. unlob does not judge whether a passage entails a claim, and it does not decide what is true. It makes independence, role, freshness and coverage explicit; deciding whether a specific sentence is supported stays with your model or your judge.

Count origins, not citations

Ask for the bar directly. ground counts independently owned origins rather than hosts, so a story republished across a syndication network counts once, and it can require a source in a given role — primary or official — before the set counts as enough.

The response says which origins it counted and which items it treated as copies, so a citation can point at the origin rather than at whichever republication ranked highest.

Two independent origins, one of them primary or official, inside a weekbash
curl -G https://api.unlob.com/ground \
  -H "x-api-key: $UNLOB_API_KEY" \
  --data-urlencode "objective=Has the regulator approved the merger?" \
  -d min_independent_origins=2 \
  -d source_roles=primary,official \
  -d max_age=7d

Let the status stop the answer

When the set does not clear the bar, ground says so — insufficient, stale or partial, with next actions naming what would change it. An agent that reads the status first can decline, ask, or search deliberately wider, instead of writing a confident answer over one source and two copies of it.

Read the opposite case carefully too: sufficient means the evidence cleared the bar you set — enough independent origins, a source in a required role, inside your freshness window — not that the claim is true. ground does not decide what is true; it reports what supports the claim and what is missing, and the judgement stays with your model.

Measure citation precision on your own tasks

Take a sample of real answers. For every citation, label it: supports the sentence, related but does not support it, contradicts it, or is a copy of another citation. Citation precision is the share that support; duplicate inflation is how many citations collapse into each origin.

When you compare retrieval layers, change only the retrieval: same model, same prompts, same task set, same grading. Otherwise a better prompt will be credited to a different search API, or the other way round.

Frequently asked questions

Does unlob check whether a passage supports a claim?

No. It does not judge whether a passage entails a claim and does not decide what is true. It reports what each passage is — source role, origin, date — which sources are independent, what was not covered, and whether the set clears the bar you set. Judging support is your model’s job, or a judge’s; the point is to hand it a set that is not structurally misleading.

Is citing a republished copy wrong?

Not wrong, but it is not additional support. Cite the origin where you can, and count the copies as one.

What is a good citation precision?

There is no universal figure; it depends on the task and on how strictly support is graded. Measure it on your own tasks before and after a retrieval change, with everything else held still.

Test it on your own tasks

Send twenty tasks and what a good result looks like, and we report what was found, what was missing, the cost, and where unlob does not fit. Or run the same evaluation yourself on 10,000 free credits a month.

API and MCP reference ↗