Get early access
For researchers

Evidence you can put in front of somebody

The fluent answer did not save you the time. It moved the time. You either spend the hours checking it or you carry the risk of not checking it.

Check one figure, so the memo survives review

A deep research answer lands in two minutes and reads well, then cannot go in front of a client until somebody walks it back to the documents. A junior analyst who takes a day and shows their working costs you less, because the checking is done by the time the memo reaches you.

40 organisations × 6 fields = 240 separate claims

What a deep research mode returns

One answer, 240 claims inside it, a list of sources at the end.

  • You can check that a source was consulted.
  • You cannot check which sentence came from which source.
  • To verify one figure you re-run that research yourself.

Cost of checking one number A search you have to do again

What comes back here

240 cells, each carrying its own source, score and verdict.

  • Each claim points at the specific document it came from.
  • Each carries a judgment on whether that document supports it.
  • Where the evidence was not there, the cell says so instead of filling.

Cost of checking one number A link on that cell

Same shape of output, same two minutes of your attention. The difference only shows up the moment somebody asks you where one of the 240 came from.

Verdict: one of four labels on a single cell, saying how well the document behind that cell supports the value in it. Confirmed, confirmed with caveats, uncertain, or mismatch. Score: how confident LoQuery is in that label, from 0 to 1. Caveat: a named fault LoQuery raises against its own answer, drawn from a fixed list.

Open the cell, so the document can answer

Suppose the sector note says Glencore paid USD 1.186 billion. The largest number in a piece is the one that gets challenged, and this is the single cell that number came out of.

Research Workbench Table View
Item
Glencore Ltd
Field
Penalty amount
Value
USD 1,186,345,850, penalty plus disgorgement, three entities collectively
Verdict
confirmed with caveats 0.78
Caveats
entity_scope_broader_than_item
Why
Score 0.78; the order names three Glencore entities collectively, so the figure is not attributable to the named item alone.
The score, the verdict, the reason and the caveat all belong to this one value, and they travel with it into the exported file.

Cell values, sources and caveats here are real and open to checking. The scores, attempt counts and source totals show the shape of a run and get replaced once a captured session record lands.

LoQuery put one document under that value, CFTC release 8534-22, and four things in that release become checkable inside a minute.

  1. Exact, and a rounding. The release reads: “Glencore is required to pay a total of $1.186 billion, which consists of the highest civil monetary penalty ($865,630,784) and highest disgorgement amount ($320,715,066) in any CFTC case.” Two kinds of money added together, so a note that calls the whole of it a fine is wrong.
  2. “Glencore” is three companies. The order settles charges against Glencore International A.G. of Switzerland, Glencore Ltd. of New York and Chemoil Corporation of New York, “collectively, Glencore”. The row asked about Glencore Ltd. One of the three is not called Glencore at all.
  3. Offset against the Justice Department. “The CFTC order recognizes and offsets certain forfeiture and penalty payments to be made to the DOJ in those cases.” Sum the two regulators' headline numbers and you have published a total that was never owed.
  4. A third authority is in the same paragraph. The UK Serious Fraud Office announced separate criminal charges, which is a boundary somebody will test if the note is scoped to US enforcement.

Every quotation above comes from the one page the cell links to, and none of it came from what the model remembers.

Research Workbench Thinking
Glencore Ltd confirmed with caveats SCORE 0.78

Evaluator Score 0.78; the order names Glencore International A.G., Glencore Ltd. and Chemoil Corporation collectively, so the figure is not attributable to the named item alone. Flagged rather than reported clean.

The evaluator writes its reason before anyone sees the score. This is that entry, unedited.

Cell values, sources and caveats here are real and open to checking. The scores, attempt counts and source totals show the shape of a run and get replaced once a captured session record lands.

USD 1.186 billion is exact, and it belongs to three companies together, only one of which the row named.

Three commodities traders, twelve cells, no clean row

LoQuery researched three enforcement actions over four fields each, and every value carries its own link, so you can settle one figure by opening one document. None came back clean.

LoQuery flagged the Freepoint Commodities name before the run started, because two separate federal actions resolve against that name in the same period, under different statutes. LoQuery researched it anyway, because you listed it.

Research Workbench Table View
Filter Confirmed clean 0 With caveats 3 Candidates 0 Filtered 0
Item name Score Verdict Why Caveats RegulatorPenalty amountDate of orderPrimary filing
Vitol Inc 2 attempts 0.79 confirmed with caveats Score 0.79; the USD 135m is a combined DOJ and Brazil resolution, and a parallel CFTC order the same day partly offsets against it. The figures are neither one number nor two that add up. parallel_action_unreported US Department of Justice justice.gov/…/vitol-inc-agrees-pay-over-135-million USD 135,000,000, combined DOJ and Brazil. A parallel CFTC order of USD 95,700,000 partly offsets against it cftc.gov/PressRoom/PressReleases/8326-20 3 December 2020 justice.gov/…/vitol-inc-agrees-pay-over-135-million Deferred prosecution agreement (FCPA) justice.gov/…/vitol-inc-agrees-pay-over-135-million
Glencore Ltd 3 attempts 0.78 confirmed with caveats Score 0.78; the order names three Glencore entities collectively, so the figure is not attributable to the named item alone. entity_scope_broader_than_item CFTC cftc.gov/PressRoom/PressReleases/8534-22 USD 1,186,345,850, penalty plus disgorgement, three entities collectively cftc.gov/PressRoom/PressReleases/8534-22 24 May 2022 cftc.gov/PressRoom/PressReleases/8534-22 CFTC order cftc.gov/PressRoom/PressReleases/8534-22
Freepoint Commodities 3 attempts 0.74 confirmed with caveats Score 0.74; a parallel CFTC order charges the same conduct under a different statute and is largely offset against this one, so neither figure alone is the answer and adding them is wrong. parallel_action_unreported US Department of Justice justice.gov/…/commodities-trading-company-98m USD 98,551,150, DOJ penalty and forfeiture. The parallel CFTC order is largely offset, with USD 7.6m disgorged cftc.gov/PressRoom/PressReleases/8834-23 14 December 2023 justice.gov/…/commodities-trading-company-98m Three-year deferred prosecution agreement, District of Connecticut justice.gov/…/commodities-trading-company-98m

Confidence chips indicate model certainty, not factual correctness. Verify critical decisions before acting on results.

Confidence 0.90 high 0.78 medium 0.60 low 0.30 caution
Export the table with its sources attached. When somebody asks where the 1.186 billion came from, the answer is a link on that cell. Their next question is whether the figure covers three entities. The cell already says so.

Cell values, sources and caveats here are real and open to checking. The scores, attempt counts and source totals show the shape of a run and get replaced once a captured session record lands.

Items researched
3
Cells, each with its own source
12
Government documents behind them
5
Rows confirmed with no caveat
0

Overlapping regulators, offset penalties

Vitol's USD 135,000,000 is a combined Justice Department and Brazil resolution. The CFTC ordered USD 95.7m the same day, in an order that offsets a portion of any criminal penalty paid to the Justice Department. Freepoint's USD 98,551,150 is the Justice Department figure alone. A second CFTC order of more than USD 91m lands against the same name under a different statute, carrying the same offset language. Either figure alone understates the resolution, and adding the two overstates it.

A tool that handed back three clean numbers here would have been confidently wrong three times, so LoQuery put a caveat on all three.

The same run, step by step →

Your browser searches, so no crawler of ours visits

LoQuery searches through a browser extension on your own machine, over your own connection, so no address of ours appears in a target firm's server logs. LoQuery buys no search, so no search supplier of ours receives a request naming who you are looking at.

Your account and your run records do live on our servers: the items, the fields, the sources and the verdicts. The searching does not. What we store →

LoQuery's reasoning runs on gpt-oss-120b, OpenAI's open-weight model, served by Groq and Cerebras. Both are American companies. No Chinese-developed model is in the path, and asked which vendors see the reasoning, that is the whole list.

Six gates run before any model

Each gate fires on the extracted values themselves rather than on the model's reasoning, so the same answer trips the same gate every time. A gate that fires lowers the score and records the reason, and LoQuery still shows the result, because deleting it would hide what the gate caught.

Gate Fires when Effect
Not-applicable relevance More than half the fields came back “not applicable” If most of what you asked for does not apply, the item is probably not the kind of thing you meant. score × 0.4
Impossible date Any date found is more than a week in the future Checked across every field, not just the ones labelled as dates. A future date in a historical answer means something was misread. score × 0.3
Temporal scope Every date found sits outside the window you asked for Six months of grace at each end, because sources round and report late. Every date being outside it is a different problem. score × 0.4
Aggregator source The entity came off a “top ten” or “list of” page A soft penalty, not a block. Sometimes a list page is genuinely the right source, and a hard rule would break those runs. confidence − 0.20
Blocked evidence chain A field could not be researched because the field it depends on was never found Distinguishes could-not-find from was-never-reachable. Punishing both the same way makes the record useless for working out what went wrong. reduced penalty
Qualifier coherence Two or more of the qualifiers you specified do not match what came back The only gate that overrides the verdict outright. One qualifier missing is a gap; two is a sign this is the wrong entity. verdict → mismatch

Read the middle column again and notice what is missing: not one of these mentions a subject. They test the shape of an answer rather than its topic, which is why nobody has to teach the system a new field before you can ask it a question about one.

When LoQuery discovers the list itself rather than taking yours, a person approves every candidate before the run spends anything researching them. Approve a discovered list yourself →

The record keeps the attempts that failed

Six months on, somebody reopens the work and asks whether LoQuery went looking for a figure it left blank, and the exported file answers that. LoQuery writes every attempt as the run happens, including the ones that returned nothing, so nobody has to reconstruct the session to check.

backend/outputs/20260728_141207_session_a4f9c1e0
  • session_batch_*.json the record. Everything below lives here
  • backend_terminal.log the full diagnostic surface
  • discovery_log.md the discovery phase, readable
task_plan
The plan the run wrote for itself. Columns, field types, what depends on what, the entity name variations it searched under, and the full configuration, so the run can be reconstructed from the file alone.
items[].final_data
Every cell, with everything attached. Value, source URL, verdict, score, caveats and reasoning, per field per item.
items[].attempts
Every attempt, including the ones that failed. Each pass with its strategy, its queries, and the named reason it stopped. This is the part that answers “did it actually look?”
intelligence_trace
How it interpreted the question. Intent analysis, field classification with whether a model or the fallback decided it, dependency structure, entity validation per wave.
search_queries_executed[]
Every query it ran. Alongside pages_fetched, so the retrieval is inspectable rather than asserted.
verdict_source
Who decided each verdict. evaluator, user or collaborative. Your corrections survive as yours, which matters when the file is reopened by somebody who was not there.

One file per session, written as the run happens rather than assembled on request. Nothing has to be exported before it exists, and nothing is discarded because the run ended badly.

Where we lose

A frontier deep-research mode LoQuery
Open-ended reasoning they win A frontier model wins, and not narrowly. Synthesising an argument, weighing competing interpretations, writing the analysis itself. Not what this is for. The architecture removes the need for cleverness at any single step, which is exactly the wrong trade if cleverness is the deliverable.
Time to first answer they win Two minutes, often less. Slower, and structurally so. Searching happens in your browser at roughly the pace a person browses, and each item gets up to three passes.
Breadth of a single question they win Will attempt anything, and produce something for all of it. Narrower. Fixed columns over a defined population. Ask it something shapeless and you get a worse answer than they will give you.
Measured accuracy they win Published evaluations. Contested, but public, and you can argue with the methodology. None to show you. Verification is manual today, because the system cannot run without a browser attached and there is nothing to point an automated harness at.
Attribution you can defend we win Sources listed for the answer. No way to tie one claim to one document without redoing the work. A source, a score and a verdict on every individual cell, plus a record of where a person overruled the machine.

Four of five go the other way, and the accuracy row is the one that costs us most to admit. If your question needs a long chain of original reasoning rather than evidence gathered and checked, use the other thing. This is built for the work where somebody is going to ask you where a number came from.

You pay what a LoQuery run costs plus a margin, with no subscription, no seat, and nothing charged for a month you did not use it. A run costs cents rather than dollars. What we charge, and why →

Get early access

Tell us the kind of research you do, and what a defensible answer has to look like before you would put it in front of somebody.

One email when it opens. No newsletter unless you ask for one.