Before you depend on it
Prove it on your workload
A search API is easy to try and hard to judge. The question that matters is whether it finds the evidence your agent needs, says when it has not, and does it at a cost that holds at your volume — and only your own tasks answer that.
Start from where you are
Which of these is you?
Each route ends at the same place — a test on your own tasks — but starts from the question you arrived with.
I am choosing a provider
Compare on the workload that matters to you, not on a headline benchmark or a price per call.
I am replacing a provider
Map your current calls, see what you lose, and run the migration against your own queries first.
My agent’s evidence is unreliable
Five articles that repeat one press release, citations that do not support the answer, an empty result read as proof of absence.
I am building a product on it
Start from the reference workload closest to yours, or embed the layer across your own customers’ agents.
I need it inside my boundary
Define the boundary, the sources, the languages and the retention first; the deployment follows from them.
Run it yourself
The evaluation, without talking to anyone
Twenty ground calls are 100 credits — at 5 credits a call, a small fraction of the 10,000 free credits a month. Score them on the three tracks below and you have the same answer we would give you.
1. Take twenty real tasks
As your agent receives them, including the awkward ones. Include some where the right answer is to abstain — that is the case most retrieval gets wrong.
2. Write down what a good result is
Before running anything: what each answer must be supported by, which source roles must be present, how fresh the evidence must be.
3. Run each through ground
Set the bar in the request — independent origins, required roles, max_age — and keep the status, the evidence and the coverage receipt for every task.
4. Score retrieval, evidence handling and the end result separately
So a good final answer cannot hide weak evidence, and a strong evidence set cannot hide a prompt problem.
Or send us the brief
What to send
The minimum we need to tell you whether unlob fits. Answer what you can; a partial brief is a fine start.
1. Representative tasks
About 20 tasks your agent actually performs, as it receives them. Sanitised or synthetic versions are fine for a first pass.
2. What a good result is
For each task, what the answer must be supported by — or that the right answer is to abstain.
3. Current approach
The provider or pipeline you use today, and what you would change about it.
4. Sources and languages
Source classes that must be present (regulators, primary filings, named publications) and the languages involved.
5. Freshness
How recent the evidence must be, and whether publication date or crawl date is the one that matters.
6. Volume and latency
Tasks per month, the peak rate, and the latency budget a call has to fit inside.
7. Deployment constraints
Where data may be processed, what may be retained, and whether the workload has to run inside your own boundary.
Answer what you can; sanitised examples are fine.
1. Representative tasks — About 20 tasks your agent actually performs, as it receives them. Sanitised or synthetic versions are fine for a first pass.
2. What a good result is — For each task, what the answer must be supported by — or that the right answer is to abstain.
3. Current approach — The provider or pipeline you use today, and what you would change about it.
4. Sources and languages — Source classes that must be present (regulators, primary filings, named publications) and the languages involved.
5. Freshness — How recent the evidence must be, and whether publication date or crawl date is the one that matters.
6. Volume and latency — Tasks per month, the peak rate, and the latency budget a call has to fit inside.
7. Deployment constraints — Where data may be processed, what may be retained, and whether the workload has to run inside your own boundary.Opens a draft with these questions already in it. A person replies.
What comes back
A report on your tasks, including the ones that went badly
Every item is reported per task. A summary that only shows the wins is an advertisement, and you would be right to discount it.
Was the required evidence found?
Per task, against the success criteria you set — not against ours.
Were the citations useful?
Whether each piece of evidence supports the claim it would be attached to, or only shares its keywords.
Were source relationships right?
Whether copies and republished stories were counted as one origin, and owners told apart from hosts.
What was missing?
The source classes, languages or periods the index did not hold, read from each coverage receipt.
What would the workload cost?
Credits per task at your volume, and the plan that volume implies, overage included.
How did it behave when it could not answer?
Whether insufficient, stale, partial and empty came back when they should have, and not when they should not.
Where does unlob not fit?
The tasks it should not be used for, said plainly. A no-fit result is a finding, not a failed sale.
How a fair test works
Three tracks, scored apart
- Retrieval
- Did it retrieve relevant, fresh enough material from the source set you require?
- Evidence handling
- Did it handle duplication, provenance, missing source classes and incomplete results correctly?
- End to end
- Did your agent produce acceptable outputs at an acceptable total cost and latency?
When comparing providers, keep the model, the prompts, the task set and the grading identical, so the only thing that differs is retrieval. Compare complete systems only when that is explicitly the question being asked.
One reading rule matters more than the rest: sufficient means the evidence cleared the bar you set — enough independent origins, a source in a required role, inside your freshness window — not that the claim is true. ground does not decide what is true; it reports what supports the claim and what is missing, and the judgement stays with your model.An evaluation scores the answer your agent gives against your criteria; the status is an input to that, not a verdict.
Our own agent benchmark runs the same method across providers on a fixed task set. Its results are pending — no figure will appear before the harness produces a run. The benchmark method.
Your tasks, and what we do with them
Tasks you send are used to run your evaluation and nothing else. They are never turned into a marketing profile, and what the API itself records about queries is described onsecurity and operations. If your tasks cannot leave your environment at all, run the evaluation yourself on your own key with the rubric above, and send us only the scores and the receipts you are comfortable sharing.
If an evaluation needs engineering time beyond running your tasks — an integration in your stack, a source class we would have to examine first — we say so and quote it before starting.
Inside your boundary
Private deployment starts with the boundary, not the API
Today unlob runs as a hosted service. Deployment inside a customer’s own boundary — the same closed product, in a national cloud, your VPC or a disconnected environment — is available by engagement and in development, and it is scoped before it is priced.
1. The boundary
National or sovereign cloud, your own VPC, or a disconnected environment — and who may operate inside it.
2. Sources and languages
Which source classes and languages the index must hold, and any you are not permitted to hold.
3. Retention and the query record
How long queries, receipts and evidence may be kept, and who holds that record.
4. Index updates
How often the index must be refreshed, and how snapshots may cross the boundary.
5. Acceptance
What the deployment has to demonstrate before it is accepted, and who signs that off.
Frequently asked questions
What does a workload evaluation cost?
Running your tasks costs nothing: 20 ground calls are 100 credits, well inside the 10,000 free credits a month, and when we run them for you we do it on our own account. If an evaluation needs engineering time beyond that — an integration in your stack, a source class we would have to examine first — we say so and quote it before starting.
Do I have to send production queries?
No. Sanitised or synthetic versions of real tasks are enough to establish whether unlob fits. What matters is that they have the shape of the real workload: the same kind of question, the same sources required, the same freshness.
What happens if unlob does not fit?
The report says so, task by task, and says why — a missing source class, a language we do not hold, a freshness window we cannot meet. That is a useful answer: it stops a bad dependency before it is built. A gap that comes up repeatedly informs what the crawler covers next, but nothing is fetched on request, so we will not promise a source we do not hold.
How is this different from the benchmark?
The benchmark runs a fixed task set through every provider under one harness, so the providers can be compared with each other. An evaluation runs your tasks against your success criteria, so you can decide whether unlob fits your workload. The first says how retrieval layers differ; the second says whether this one is right for you.
Start with twenty tasks
Send the brief and a person replies, or take a free key and run the same evaluation yourself.