Evidence you can put in front of somebody
The fluent answer did not save you the time. It moved the time. You either spend the hours checking it or you carry the risk of not checking it.
Check one figure, so the memo survives review
A deep research answer is the intern who is back in two minutes with nothing to show how they got there. It reads well, and it cannot go in front of a client until somebody walks it back to the documents. A junior analyst who takes a day and shows their working costs you less, because the checking is done by the time the memo reaches you.
40 organisations × 6 fields = 240 separate claims
What a deep research mode returns
One answer, 240 claims inside it, a citation under each sentence.
- Each citation shows which page a sentence came from.
- None of them says how well that page supports the sentence.
- To settle one figure you open the page and re-read it yourself.
Cost of checking one number A document you have to re-read
What comes back here
240 cells, each carrying its own source, score and verdict.
- Each claim points at the specific document it came from.
- Each carries a judgment on whether that document supports it.
- Where the evidence was not there, the cell says so instead of filling in a value.
Cost of checking one number A link on that cell
Same shape of output, same two minutes of your attention. The difference only shows up the moment somebody asks you where one of the 240 came from.
Verdict: one of four labels on a single cell, saying how well the document behind that cell supports the value in it. Confirmed, confirmed with caveats, uncertain, or mismatch. Score: how confident LoQuery is in that label, from 0 to 1. Caveat: a named fault LoQuery raises against its own answer, drawn from a fixed list.
Open the cell, so the document can answer
Suppose the sector note says Glencore paid USD 1.186 billion. The largest number in a piece is the one that gets challenged, and this is the single cell that number came out of.
- Item
- Glencore Ltd
- Field
- Penalty amount
- Value
- USD 1,186,345,850, penalty plus disgorgement, three entities collectively
- Verdict
- confirmed with caveats 0.78
- Caveats
- entity_scope_broader_than_item
- Why
- Score 0.78; the order settles with three entities collectively defined as Glencore, so the figure is not attributable to the named item alone.
Cell values, sources and caveats here are real and open to checking. The scores, attempt counts and source totals show the shape of a run and get replaced once a captured session record lands.
LoQuery put one document under that value, CFTC release 8534-22, and four things in that release become checkable inside a minute.
- Exact, and a rounding. The release reads: “Glencore is required to pay a total of $1.186 billion, which consists of the highest civil monetary penalty ($865,630,784) and highest disgorgement amount ($320,715,066) in any CFTC case.” Two kinds of money added together, so a note that calls the whole of it a fine is wrong.
- “Glencore” is three companies. The order settles charges against Glencore International A.G. of Switzerland, Glencore Ltd. of New York and Chemoil Corporation of New York, “collectively, Glencore”. The row asked about Glencore Ltd. One of the three is not called Glencore at all.
- Offset against the Justice Department. “The CFTC order recognizes and offsets certain forfeiture and penalty payments to be made to the DOJ in those cases.” Sum the two regulators' headline numbers and you have published a total that was never owed.
- A third authority is in the same paragraph. The UK Serious Fraud Office announced separate criminal charges, which is a boundary somebody will test if the note is scoped to US enforcement.
Every quotation above comes from the one page the cell links to, and none of it was recalled from training.
USD 1.186 billion is exact, and it belongs to three companies together, only one of which the row named.
Three commodities traders, twelve cells, no clean row
LoQuery researched three enforcement actions over four fields each, and every value carries its own link, so you can settle one figure by opening one document. None came back clean, and on this set none of them should have.
LoQuery flagged the Freepoint Commodities name before the run started, because two separate federal actions resolve against that name in the same period, under different statutes. LoQuery researched it anyway, because you listed it.
| Item name | Score | Verdict | Why | Caveats | Regulator | Penalty amount | Date of order | Primary filing |
|---|---|---|---|---|---|---|---|---|
| Vitol Inc 2 attempts | 0.79 | confirmed with caveats | Score 0.79; the USD 135m is a combined DOJ and Brazil resolution, and a parallel CFTC order the same day partly offsets against it. The figures are neither one number nor two that add up. | parallel_action_unreported | US Department of Justice justice.gov/…/vitol-inc-agrees-pay-over-135-million | USD 135,000,000, a combined DOJ and Brazil resolution justice.gov/…/vitol-inc-agrees-pay-over-135-million | 3 December 2020 justice.gov/…/vitol-inc-agrees-pay-over-135-million | Deferred prosecution agreement (FCPA) justice.gov/…/vitol-inc-agrees-pay-over-135-million |
| Glencore Ltd 3 attempts | 0.78 | confirmed with caveats | Score 0.78; the order settles with three entities collectively defined as Glencore, so the figure is not attributable to the named item alone. | entity_scope_broader_than_item | CFTC cftc.gov/PressRoom/PressReleases/8534-22 | USD 1,186,345,850, penalty plus disgorgement, three entities collectively cftc.gov/PressRoom/PressReleases/8534-22 | 24 May 2022 cftc.gov/PressRoom/PressReleases/8534-22 | CFTC order cftc.gov/PressRoom/PressReleases/8534-22 |
| Freepoint Commodities 3 attempts | 0.74 | confirmed with caveats | Score 0.74; a parallel CFTC order charges the same conduct under a different statute and is largely offset against this one, so neither figure alone is the answer and adding them is wrong. | parallel_action_unreported | US Department of Justice justice.gov/…/commodities-trading-company-98m | USD 98,551,150, a DOJ penalty and forfeiture justice.gov/…/commodities-trading-company-98m | 14 December 2023 justice.gov/…/commodities-trading-company-98m | Three-year deferred prosecution agreement, District of Connecticut justice.gov/…/commodities-trading-company-98m |
| Item name | Score | Verdict | Why | Caveats | Regulator | Penalty amount | Date of order | Primary filing |
|---|---|---|---|---|---|---|---|---|
| Vitol Inc 2 attempts | 0.79 | confirmed with caveats | Score 0.79; the USD 135m is a combined DOJ and Brazil resolution, and a parallel CFTC order the same day partly offsets against it. The figures are neither one number nor two that add up. | parallel_action_unreported | US Department of Justice justice.gov/…/vitol-inc-agrees-pay-over-135-million | USD 135,000,000, a combined DOJ and Brazil resolution justice.gov/…/vitol-inc-agrees-pay-over-135-million | 3 December 2020 justice.gov/…/vitol-inc-agrees-pay-over-135-million | Deferred prosecution agreement (FCPA) justice.gov/…/vitol-inc-agrees-pay-over-135-million |
| Glencore Ltd 3 attempts | 0.78 | confirmed with caveats | Score 0.78; the order settles with three entities collectively defined as Glencore, so the figure is not attributable to the named item alone. | entity_scope_broader_than_item | CFTC cftc.gov/PressRoom/PressReleases/8534-22 | USD 1,186,345,850, penalty plus disgorgement, three entities collectively cftc.gov/PressRoom/PressReleases/8534-22 | 24 May 2022 cftc.gov/PressRoom/PressReleases/8534-22 | CFTC order cftc.gov/PressRoom/PressReleases/8534-22 |
| Freepoint Commodities 3 attempts | 0.74 | confirmed with caveats | Score 0.74; a parallel CFTC order charges the same conduct under a different statute and is largely offset against this one, so neither figure alone is the answer and adding them is wrong. | parallel_action_unreported | US Department of Justice justice.gov/…/commodities-trading-company-98m | USD 98,551,150, a DOJ penalty and forfeiture justice.gov/…/commodities-trading-company-98m | 14 December 2023 justice.gov/…/commodities-trading-company-98m | Three-year deferred prosecution agreement, District of Connecticut justice.gov/…/commodities-trading-company-98m |
Press Esc or click outside to close
Confidence chips indicate model certainty, not factual correctness. Verify critical decisions before acting on results.
Sources here are real; scores and attempt counts are illustrative until a session capture lands.
- Items researched
- 3
- Cells, each with its own source
- 12
- Government documents behind them
- 3
- Rows confirmed with no caveat
- 0
Every document behind that run is a government filing, which is the easy case. The student run on this site has the hard case, a 1953 incident where LoQuery located no government or military file, and its empty Primary source cell makes the same argument from the other direction. That run, and the cell it turns on →
Nobody wrote a module for your desk
A market research desk wants pricing, headcount and funding for every company on a list. An enforcement desk wants the order, the penalty and the filing. Same table, same three labels.
Both desks make the same choice first. You paste the population, or you describe it and LoQuery finds it before anything is researched.
Direct
Vitol Inc, Glencore Ltd, Freepoint Commodities
It researches each item against your criteria and returns one row per item.
- You paste the list
- It researches each item
- Table with a source on every cell
Use it when You already know the population and need evidence about it.
Discovery
Aerial incidents over Canada or the Great Lakes that a government or military body formally investigated
It finds the population first, in waves, then hands the list back for you to approve before it spends anything researching it. Set a discovery budget and it can run straight through.
- You describe the category
- It collects candidates in waves
- It validates and removes duplicates
- You approve, strike or send it back
- Then it researches the approved list
Use it when You do not know who is in the set, and the answer depends on getting that right.
A model writes each run's configuration from your description, and the enforcement run above used exactly that. Override the fields you have an opinion about, so a desk nobody built for still reaches the sources it trusts.
Auto-detect. Describe the question in plain language and a model writes the whole configuration for this run. Most people never leave here.
Build Your Own Type. You supply the configuration. Whatever you set is locked and never overwritten.
There is no third path, and specifically no library of prebuilt domain templates. Six of those existed once and were deleted, because a template list is where a general system quietly becomes a specialist in whatever is on the list.
-
entity_typeWhat counts as one row clinical_trials -
discovery_preferred_sourcesSpecialist sources to reach for first clinicaltrials.gov, ema.europa.eu, who.int -
tier_overridesPromote or demote specific domains for this run clinicaltrials.gov: 0.95 -
entity_confidence_thresholdHow sure it has to be before a candidate survives 0.70 -
additional_field_typesField names it would not recognise auto -
dedup_key_patternWhat makes two names the same entity auto -
task_doer_contextExtraction rules for this subject auto -
field_guidance_overridesWhere to look for a novel field type auto
Field names are the product's own. The values beside them are illustrative of a subject nobody configured, rather than a captured run.
Blanks are not defaults, they are deferrals. Anything you leave empty is generated for this run, and if that generation fails the static foundation underneath still answers, with a more generic result, not a broken one.
Give the run what only you know
Item warnings · 1
Freepoint Commodities: Entity screening flagged 'Freepoint Commodities' (Two separate federal actions resolve against this name in the same period, under different statutes.) Researching it anyway because you listed it explicitly. Treat its results with this in mind.
Knowledge panel
-
Entity definitions wave boundary
No entity definitions added yet.
e.g. 'ABC Corp also operates as XYZ Holdings'
-
Validation rules verdict boundary
No validation rules added yet.
e.g. 'Settlement must be >$1M'
-
Knowledge triples wave boundary · verdict boundary
No knowledge triples added yet.
Subject · Predicate · Object
-
Source authority wave boundary
No source authority overrides added yet.
Domain pattern (e.g. sec.gov)
Freepoint's USD 98,551,150 is the Justice Department figure alone. A second CFTC order of more than USD 91m lands against the same name under a different statute, carrying the same offset language. Either figure alone understates the resolution, and adding the two overstates it.
You know something the model does not, and that is what settles it. Type it in and LoQuery holds the rest of the run to it: what counts as this entity, what an answer must pass, which domain outranks which.
Your rule holds from the next item onward, and the rows judged before it stand as LoQuery judged them.
Your browser searches today, so no crawler of ours visits
LoQuery searches through a browser extension on your own machine, over your own connection. Today every search request leaves your machine rather than ours, so no address of ours appears in a target firm's server logs and no search service ties your targets to us.
Your account and your run records do live on our servers: the items, the fields, the sources and the verdicts. Today the searching does not. What we store →
LoQuery's reasoning runs on gpt-oss-120b and gpt-oss-20b, OpenAI's open-weight models, served by Groq and Cerebras. Both are American companies. No Chinese-developed model is in the path, and if you ask which vendors see the reasoning, that is the whole list.
Six gates run before any model judges
Each gate fires on the extracted values themselves rather than on the model's reasoning, so the same answer trips the same gate every time. A gate that fires lowers the score and records the reason, and LoQuery still shows the result, because deleting it would hide what the gate caught.
Evaluator Score 0.78; the order names Glencore International A.G., Glencore Ltd. and Chemoil Corporation collectively, so the figure is not attributable to the named item alone. Flagged rather than reported clean.
Sources here are real; scores and attempt counts are illustrative until a session capture lands.
| Gate | Fires when | Effect |
|---|---|---|
| Not-applicable relevance | More than half the fields came back “not applicable” If most of what you asked for does not apply, the item is probably not the kind of thing you meant. | score × 0.4 |
| Impossible date | Any date found is more than a week in the future Checked across every field, not just the ones labelled as dates. A future date in a historical answer means something was misread. | score × 0.3 |
| Temporal scope | Every date found sits outside the window you asked for Six months of grace at each end, because sources round and report late. Every date being outside it is a different problem. | score × 0.4 |
| Aggregator source | The entity came off a “top ten” or “list of” page A soft penalty rather than a block. Sometimes a list page is genuinely the right source, and a hard rule would break those runs. | confidence − 0.20 |
| Blocked evidence chain | A field could not be researched because the field it depends on was never found Distinguishes could-not-find from was-never-reachable. Punishing both the same way makes the record useless for working out what went wrong. | reduced penalty |
| Qualifier coherence | Two or more of the qualifiers you specified do not match what came back The only gate that overrides the verdict outright. One qualifier missing is a gap; two is a sign this is the wrong entity. | verdict → mismatch |
Read the middle column again and notice what is missing: not one of these mentions a subject. They test the shape of an answer rather than its topic, which is why nobody has to teach the system a new field before you can ask it a question about one.
The record keeps the attempts that failed
Six months on, somebody reopens the work and asks whether LoQuery went looking for a figure it left blank, and the exported file answers that. LoQuery writes every attempt as the run happens, including the ones that returned nothing, so nobody has to reconstruct the session to check. Nothing carries over from an earlier run either. A verdict cache was proposed as a latency fix and refused, so a figure you check in March was judged against March's web. Every verdict in the file also names whose judgment produced it: the machine's, yours, or both. A deep research answer has no field for that, because nothing in it expects a person to overrule it.
- Attempt 1 Precision up to 6 queries
Narrow, well-targeted searches aimed at the sources most likely to be authoritative, and it only keeps sources that clear a high quality bar.
- Attempt 2 Breadth up to 12 queries
Wider phrasing and more sources, informed by what the first attempt failed to find, with the source bar lowered to match.
- Attempt 3 Exhaustive up to 18 queries
The widest search and the lowest source bar. It takes the best available and reports honestly on how good that was.
Then all three are merged. The final row is built from the best data any attempt produced, not the last one to run. An attempt that turned out to be about the wrong entity is excluded from the merge rather than averaged in.
It stops early when a further attempt would not learn anything
-
PLATEAU_EARLY_EXIT - The score barely moved and the next attempt kept finding the same pages. Both conditions have to hold. A small score gain on genuinely new sources is progress; the same sources again is not.
-
EARLY_EXIT_BARREN - Nothing was found for the dependent fields after the second attempt. A third pass over an empty result is a way of spending money to reach the same conclusion.
-
NEW7_BREADTH_PIVOT - Barren, but the item's identity anchor was found. The opposite call. Knowing which entity this is means a wider third attempt has something to work from, so it runs.
Every exit is written into the run record by name, so a run that finished in one attempt can be told apart from a run that gave up.
session_batch_*.jsonthe record. Everything below lives herebackend_terminal.logthe full diagnostic surfacediscovery_log.mdthe discovery phase, readable
task_plan- The plan the run wrote for itself. Columns, field types, what depends on what, the entity name variations it searched under, and the full configuration, so the run can be reconstructed from the file alone.
items[].final_data- Every cell, with everything attached. Value, source URL, verdict, score, caveats and reasoning, per field per item.
items[].attempts- Every attempt, including the ones that failed. Each pass with its strategy, its queries, and the named reason it stopped. This is the part that answers “did it actually look?”
intelligence_trace- How it interpreted the question. Intent analysis, field classification with whether a model or the fallback decided it, dependency structure, entity validation per wave.
search_queries_executed[]- Every query it ran. Alongside pages_fetched, so the retrieval is inspectable rather than asserted.
verdict_source- Who decided each verdict. evaluator, user or collaborative. Your corrections survive as yours, which matters when the file is reopened by somebody who was not there.
One file per session, written as the run happens rather than assembled on request. Nothing has to be exported before it exists, and nothing is discarded because the run ended badly.
Where we lose
One boundary, stated before you test it. LoQuery does regulatory and enforcement research: agency releases, published orders, filings, the open record. It is not a legal research tool, because it reaches no subscription database and no case-law citator. Ask it for precedent and it will search the open record anyway, because you asked.
| A frontier deep-research mode | LoQuery | |
|---|---|---|
| Open-ended reasoning they win | A frontier model wins, and not narrowly. Synthesising an argument, weighing competing interpretations, writing the analysis itself. | Not what this is for. The pipeline is arranged so no single step has to be clever, which is exactly the wrong trade if cleverness is the deliverable. |
| Time to first answer they win | Two minutes, often less. | Slower, and structurally so: each item gets up to three passes, and searching happens in your browser today at roughly the pace a person browses. |
| Breadth of a single question they win | Will attempt anything, and produce something for all of it. | Narrower. Fixed columns over a defined population. Ask it something shapeless and you get a worse answer than they will give you. |
| Measured accuracy they win | Published evaluations. Contested, but public, and you can argue with the methodology. | None to show you. Verification is manual today, because every run needs a browser attached and there is nothing yet to point an automated harness at. |
| Attribution you can defend we win | A citation under each sentence, and no record of whether that page supports it. | A source, a score and a verdict on every individual cell, plus a record of where a person overruled the machine. |
Four of five go the other way, and the accuracy row is the one that costs us most to admit. If your question needs a long chain of original reasoning rather than evidence gathered and checked, use the other thing. This is built for the work where somebody is going to ask you where a number came from.
You pay what a LoQuery run costs plus a margin, with no subscription, no seat, and nothing charged for a month you did not use it. Cents a run rather than dollars, and your margin is part of what holds school pricing at cost. A tool that earns by keeping you engaged cannot afford a cell that reads uncertain. Every LoQuery run pays for itself, so nothing in our economics needs you to ask one more question. What we charge, and why →