Get early access

What to assign when the detector is off

64.7% of Canadian teachers surveyed in grades 6-12 report no training on student AI use. A verification pass students hand in, and the rubric that marks it.

You cannot tell who wrote it. Teachers keep saying some version of that sentence in public, and the tools sold to settle it do not work well enough to bet a student’s record on.

64.7% of surveyed Canadian teachers in grades 6 to 12 report no training on identifying student AI use, and only 42.3% say their school has a policy on it at all (Fraser Institute survey, run by Leger). So this post skips the detection argument almost entirely. It offers the assignment to set instead: hand students a machine-produced research table, make them open the sources, and mark the verdicts they change. It fits one page and one period, and it marks on four bands.

One honesty note before anything else: this post reports no classroom result, because we have not run one. Both figures above are teacher self-reports from a survey rather than an audit of school policy files. And the reason we are writing it now is on our own weakness register, which says the product’s four-word verdict vocabulary is opaque to a fifteen-year-old. Teaching that vocabulary here is the cheapest fix we have.

Vanderbilt turned its AI detector off in 2023

Vanderbilt University switched off the AI-writing detection built into its learning platform in 2023, and the reasoning was arithmetic rather than ideology. Of roughly 75,000 papers submitted in 2022, about 750 could have been falsely flagged. Each one of those is a meeting, an accusation, and a student who did the work sitting across from someone who doubts it.

Stanford researchers ran 91 TOEFL essays through the same class of tool. Every essay was written by a non-native English speaker sitting a proficiency exam. More than half came back labelled AI-generated, and 89 of the 91 were flagged by at least one of the seven detectors tested.

Why AI detectors flag writing a person wrote

A flag arrives without a reason attached. What the Stanford result shows is that the flag tracked something other than authorship. All 91 essays had human authors, and what those authors had in common was writing English as a second language. A teacher holding a flag is holding a score about prose, not a finding about who wrote the page.

The detection case ends there. The assessment question does not.

Assign the checking, not the text

Grade the work a machine cannot do for a student: opening the source, reading what is actually there, and coming back with a judgment they will defend out loud. The artifact a machine can produce stops being the thing you mark.

This is worth saying plainly, because it is the line we care most about: it does not put more AI in front of a class. It puts one machine’s output in front of them so they can take it apart.

64.7% of surveyed Canadian teachers in grades 6 to 12 report no training on identifying student AI use.

Call it AI literacy on the unit plan, because that is what it is. Canada sits 44th of 47 countries on AI training and literacy in KPMG’s 2025 study with the University of Melbourne, and 42nd on trust. The habit this assignment builds is the one that number says is missing.

Five things change in the assignment, and only five:

The task
The assignment now Research three enforcement actions and write a summary of each.
The assignment instead Here is the machine's table for those three. Find where it is weakest and show your working.
What the student does
The assignment now Prompts a chatbot, reads a fluent paragraph, rewrites it in their own words.
The assignment instead Opens sources. Compares what the cell claims against what the document says. Overrides what does not hold.
What they hand in
The assignment now Prose you cannot attribute.
The assignment instead A record of which claims they checked, which they changed, and the reasoning they gave for each.
What you are grading
The assignment now Whether the writing is good, and whether you believe they wrote it.
The assignment instead Judgment. Whether they went to the source, and whether their disagreement was well founded.
If they use AI
The assignment now The problem.
The assignment instead The point. The machine is supposed to do the collecting.

Three rows from a Grade 11 run, none of them clean

The run below comes from a Grade 11 unit on evidence and extraordinary claims. The question was which aerial incidents over Canada or the Great Lakes a government or military body formally investigated. Three items went in and no row came back clean.

Every cell holds a value, a link, and a verdict: one word from a fixed set of four, saying how well that document supports the value. Some cells also carry a caveat, a coded reason to read the value with care.

Grade 11 run · 3 items · 4 fields
fields   date
         official explanation
         primary source
         still unexplained

Shag Harbour 1967
  score    0.82
  verdict  confirmed with caveats
  filled   4 of 4
  caveats
    uncalibrated_authority_domain

Kinross 1953
  score    0.71
  verdict  confirmed with caveats
  filled   4 of 4
  caveats
    source_contested
    uncalibrated_authority_domain

Lake Michigan 1994
  score    0.18
  verdict  mismatch
  filled   0 of 4

Shag Harbour, 4 October 1967. The RCMP and Canadian Forces investigated and reached no determination. The primary source cell points at a Library and Archives Canada research list. Its archival items are National Research Council records, and the question asked for the RCMP or DND file. That gap is what the caveat on the row names.

Kinross, 23 November 1953. Ask that row for its primary source and the answer is none located. The best document the run found is a user-submitted database entry, and a wiki entry is not the USAF report it stands in for. A USAF determination does exist: an F-89C tracking an off-course RCAF C-47, lost over Lake Superior. The run located no government or military file, and the RCAF pilot named in the determination denied being intercepted.

Lake Michigan 1994. The 1994 reports are a cluster of sightings across several towns, and no single incident record resolves. The machine left the row empty and said why.

The cell values and sources in that run are real and open to checking. The scores and fill counts show the shape of a run and will be replaced once a captured session record lands.

A student who can explain why nothing goes in that third row has done the thinking the unit was for. Kinross sets the same problem in a harder form. It has an explanation on record and no document behind it, and only the primary source cell says so. Mark the student who notices that gap.

How to run the verification pass in one period

  1. Hand out a table where every cell carries the document it came from.
  2. Name the cells that must be checked, or let each student pick their own three.
  3. Give out the sheet below, and one period to fill it.
  4. Collect it and mark it against the four bands further down.
VERIFICATION PASS
one page, one period

For every cell you checked:
  1. What the document says, in
     its own words.
  2. Your verdict: confirmed /
     confirmed with caveats /
     uncertain / mismatch.
  3. If yours differs from the
     machine's, the line that
     forced the change.

Once, at the bottom:
  4. One cell you could not
     settle, and what would
     settle it.

Nothing on that page can be produced without opening a source. A pass written from the table alone is visible on sight, because it has no quotations in it.

We do not know how long this takes with thirty students in a room. One period is an estimate from the shape of the work rather than a timing from a classroom, and we would honestly rather publish that sentence than guess.

Four verdict words, and the rubric that marks them

The vocabulary is closed on purpose. Choosing one of four words asks what kind of answer the student is holding, which is a harder and more useful question than “is this right?”.

confirmed

We found it, and the source backs it up.

What it actually means

The evidence cleared the threshold and the source is one the system rates as authoritative for this kind of question. It still means go and look. A confirmed cell is a claim with a receipt attached; whether the receipt holds is your check to make.

Ask a student Open the source. Does it say what the cell says?

confirmed with caveats

We found it, but there is something you should know first.

What it actually means

The claim is supported and something about it needs qualifying: the document covers more than the thing you asked about, the source is weaker than we would like, or a date sits near the edge of the window you set. The caveat is written out with a code, so it is a category rather than a mood.

Ask a student Read the caveat first, then decide whether the answer still works for your question.

uncertain

We could not establish it, so we are not going to guess.

What it actually means

The system searched, read pages, and did not find evidence it was willing to stand behind. It scores itself zero and says so. This is the hardest behaviour to build and the easiest to skip, because a system that always produces something looks better and helps less.

Ask a student Ask why. Is it not on the open web, or was the question not answerable as asked?

mismatch

We found something, but it is about the wrong thing.

What it actually means

What came back does not resolve to the entity you named. A different company with a similar name, a cluster of events where you asked about one, a person rather than an incident. The result is kept and flagged rather than quietly dropped, because knowing the search went sideways is more useful than an empty cell.

Ask a student Look at what it did find. Usually it tells you the question needs to be more specific.

And on every verdict, a record of who decided it

  • evaluator the machine reached this on its own
  • user a person overruled the machine and this is their call
  • collaborative the machine's verdict stood, with a person's reasoning attached

That field is the assessable artifact. It is the difference between a student who accepted everything and a student who argued with three cells and wrote down why.

Confirmed with caveats is the one worth teaching first, because it is the state most real evidence is in. A cell reporting USD 1,186,345,850 against Glencore carries a caveat saying the CFTC order names Glencore International A.G., Glencore Ltd. and Chemoil Corporation collectively. The figure is correct. It just does not belong to the one company the question named.

Mark each pass on four bands

  1. Not yet. Every verdict accepted. No source opened, or opened with nothing written down.
  2. Developing. Sources opened, all four verdicts agreed with, reasons that restate the cell in other words.
  3. Proficient. At least one verdict changed, with the specific line in the document that forced the change.
  4. Strong. Uncertain used where the evidence ran out, with a statement of what would settle it.

Prose quality is not marked here, and neither is your confidence about who typed the original text. One of those was never the learning objective, and the other is not recoverable.

What we cannot claim yet

Two limits belong to the assignment itself. The pass needs a table where each cell carries its own document, so a tool that returns prose with a bibliography at the end will not support it. And nothing here assesses writing. If your unit is graded on composition, this replaces the wrong half of it.

Set this on Monday with any table that carries its sources

You do not need us for this, and we mean that literally. Take any research table where every cell carries its own document, hand out the four prompts above, and mark what comes back on the four bands. The vocabulary is free to lift.

The signal it worked is narrow. At least one student changed a verdict and can point at the line that made them. If every pass comes back with every verdict agreed with, the table was too clean.

What a teacher asks first

Isn't this just more AI in the classroom?

It is one machine output, brought in specifically to be taken apart. Students do not converse with anything; they open documents and argue with a table. The skill being marked is checking, which is the opposite of accepting.

What stops a student inventing the quotations?

Every cell links to its document, so a spot-check takes seconds: open the link, search the quoted words. An invented quotation is a faster catch than an AI-written essay ever was.

Do I need LoQuery to run this?

No. Any research table where every value carries its own source supports the pass, and the rubric is free to use. What our tool adds is the table itself: verdicts, scores, and caveats generated per cell, including rows it refuses to fill.

How can I tell if a student used AI?

Reliably, you cannot, and that is the premise of this post rather than a gap in it. What you can see is whether a student opened the sources. The verification pass asks for what a document says in its own words, and a quotation that is not in the document it names takes seconds to catch.

The educators’ page carries the longer argument and an honest list of what is not built. What we can offer now is a place in the first group of classrooms, and some influence over what the teacher’s view looks like when it is built. Join the first classrooms, and tell us what you teach and what a research unit looks like for you.