Skip to content
unlob

Before you depend on it

Prove it on your workload

A search API is easy to try and hard to judge. The question that matters is whether it finds the evidence your agent needs, says when it has not, and does it at a cost that holds at your volume — and only your own tasks answer that.

Start from where you are

Which of these is you?

Each route ends at the same place — a test on your own tasks — but starts from the question you arrived with.

Run it yourself

The evaluation, without talking to anyone

Twenty ground calls are 100 credits — at 5 credits a call, a small fraction of the 10,000 free credits a month. Score them on the three tracks below and you have the same answer we would give you.

  1. 1. Take twenty real tasks

    As your agent receives them, including the awkward ones. Include some where the right answer is to abstain — that is the case most retrieval gets wrong.

  2. 2. Write down what a good result is

    Before running anything: what each answer must be supported by, which source roles must be present, how fresh the evidence must be.

  3. 3. Run each through ground

    Set the bar in the request — independent origins, required roles, max_age — and keep the status, the evidence and the coverage receipt for every task.

  4. 4. Score retrieval, evidence handling and the end result separately

    So a good final answer cannot hide weak evidence, and a strong evidence set cannot hide a prompt problem.

Or send us the brief

What to send

The minimum we need to tell you whether unlob fits. Answer what you can; a partial brief is a fine start.

  1. 1. Representative tasks

    About 20 tasks your agent actually performs, as it receives them. Sanitised or synthetic versions are fine for a first pass.

  2. 2. What a good result is

    For each task, what the answer must be supported by — or that the right answer is to abstain.

  3. 3. Current approach

    The provider or pipeline you use today, and what you would change about it.

  4. 4. Sources and languages

    Source classes that must be present (regulators, primary filings, named publications) and the languages involved.

  5. 5. Freshness

    How recent the evidence must be, and whether publication date or crawl date is the one that matters.

  6. 6. Volume and latency

    Tasks per month, the peak rate, and the latency budget a call has to fit inside.

  7. 7. Deployment constraints

    Where data may be processed, what may be retained, and whether the workload has to run inside your own boundary.

The brief, as plain texttext
Answer what you can; sanitised examples are fine.

1. Representative tasks — About 20 tasks your agent actually performs, as it receives them. Sanitised or synthetic versions are fine for a first pass.


2. What a good result is — For each task, what the answer must be supported by — or that the right answer is to abstain.


3. Current approach — The provider or pipeline you use today, and what you would change about it.


4. Sources and languages — Source classes that must be present (regulators, primary filings, named publications) and the languages involved.


5. Freshness — How recent the evidence must be, and whether publication date or crawl date is the one that matters.


6. Volume and latency — Tasks per month, the peak rate, and the latency budget a call has to fit inside.


7. Deployment constraints — Where data may be processed, what may be retained, and whether the workload has to run inside your own boundary.
Email the brief

Opens a draft with these questions already in it. A person replies.

What comes back

A report on your tasks, including the ones that went badly

Every item is reported per task. A summary that only shows the wins is an advertisement, and you would be right to discount it.

  • Was the required evidence found?

    Per task, against the success criteria you set — not against ours.

  • Were the citations useful?

    Whether each piece of evidence supports the claim it would be attached to, or only shares its keywords.

  • Were source relationships right?

    Whether copies and republished stories were counted as one origin, and owners told apart from hosts.

  • What was missing?

    The source classes, languages or periods the index did not hold, read from each coverage receipt.

  • What would the workload cost?

    Credits per task at your volume, and the plan that volume implies, overage included.

  • How did it behave when it could not answer?

    Whether insufficient, stale, partial and empty came back when they should have, and not when they should not.

  • Where does unlob not fit?

    The tasks it should not be used for, said plainly. A no-fit result is a finding, not a failed sale.

How a fair test works

Three tracks, scored apart

Retrieval
Did it retrieve relevant, fresh enough material from the source set you require?
Evidence handling
Did it handle duplication, provenance, missing source classes and incomplete results correctly?
End to end
Did your agent produce acceptable outputs at an acceptable total cost and latency?

When comparing providers, keep the model, the prompts, the task set and the grading identical, so the only thing that differs is retrieval. Compare complete systems only when that is explicitly the question being asked.

One reading rule matters more than the rest: sufficient means the evidence cleared the bar you set — enough independent origins, a source in a required role, inside your freshness window — not that the claim is true. ground does not decide what is true; it reports what supports the claim and what is missing, and the judgement stays with your model.An evaluation scores the answer your agent gives against your criteria; the status is an input to that, not a verdict.

Our own agent benchmark runs the same method across providers on a fixed task set. Its results are pending — no figure will appear before the harness produces a run. The benchmark method.

Your tasks, and what we do with them

Tasks you send are used to run your evaluation and nothing else. They are never turned into a marketing profile, and what the API itself records about queries is described onsecurity and operations. If your tasks cannot leave your environment at all, run the evaluation yourself on your own key with the rubric above, and send us only the scores and the receipts you are comfortable sharing.

If an evaluation needs engineering time beyond running your tasks — an integration in your stack, a source class we would have to examine first — we say so and quote it before starting.

Inside your boundary

Private deployment starts with the boundary, not the API

Today unlob runs as a hosted service. Deployment inside a customer’s own boundary — the same closed product, in a national cloud, your VPC or a disconnected environment — is available by engagement and in development, and it is scoped before it is priced.

  1. 1. The boundary

    National or sovereign cloud, your own VPC, or a disconnected environment — and who may operate inside it.

  2. 2. Sources and languages

    Which source classes and languages the index must hold, and any you are not permitted to hold.

  3. 3. Retention and the query record

    How long queries, receipts and evidence may be kept, and who holds that record.

  4. 4. Index updates

    How often the index must be refreshed, and how snapshots may cross the boundary.

  5. 5. Acceptance

    What the deployment has to demonstrate before it is accepted, and who signs that off.

Frequently asked questions

What does a workload evaluation cost?

Running your tasks costs nothing: 20 ground calls are 100 credits, well inside the 10,000 free credits a month, and when we run them for you we do it on our own account. If an evaluation needs engineering time beyond that — an integration in your stack, a source class we would have to examine first — we say so and quote it before starting.

Do I have to send production queries?

No. Sanitised or synthetic versions of real tasks are enough to establish whether unlob fits. What matters is that they have the shape of the real workload: the same kind of question, the same sources required, the same freshness.

What happens if unlob does not fit?

The report says so, task by task, and says why — a missing source class, a language we do not hold, a freshness window we cannot meet. That is a useful answer: it stops a bad dependency before it is built. A gap that comes up repeatedly informs what the crawler covers next, but nothing is fetched on request, so we will not promise a source we do not hold.

How is this different from the benchmark?

The benchmark runs a fixed task set through every provider under one harness, so the providers can be compared with each other. An evaluation runs your tasks against your success criteria, so you can decide whether unlob fits your workload. The first says how retrieval layers differ; the second says whether this one is right for you.

Start with twenty tasks

Send the brief and a person replies, or take a free key and run the same evaluation yourself.

API and MCP reference ↗