How this is made

Editorial policy

How this record is made, what gets verified, and what we will not write.
Recorded on
2026-08-24

The size of this record right now

267 problems · 4,405 evidence rows · 3,529 of them carry a source you can open · 1,212 distinct sources.

The remaining 876 rows are not blank by neglect but blank by decision. Where we could not reach a source we do not delete the row; we leave the cell empty and write down why — "the original PDF exceeded the size we could open," for instance. Had we deleted them, this ratio would read 100% and mean nothing.

These numbers are counted from the documents as they stand. We do not use the counts a document declares about itself.

Every problem is written in the same thirteen parts

“What is happening?” · “Whose problem is this?” · “Where does this problem end?” · “What is the state now, and what should it be?” · “How big is it?” · “Under what conditions does it arise?” · “What has been tried?” · “What was found?” · “Why is it still unsolved?” · “What observation would mean it is solved?” · “What is it connected to?” · “What these sources do not say” · “See the evidence”.

Fixing the shape is a discipline, not a convenience. When the parts are set, a part we could not answer stays visibly unanswered. Free-form writing lets the unknown disappear quietly by simply not being written; thirteen fixed parts do not allow that.

The last part (“See the evidence”) is a table. Every claim in the body should be traceable to a row in it.

What counts as a problem here

The starting point is the record public bodies keep about themselves — audit reports, legislative records, official statistics, regulators' notices and determinations, published plans. What this record does is gather what those documents already said and set it down in one shape.

Six countries are covered — Korea, the United States, Australia, the United Kingdom, New Zealand and Canada — and problems belonging to no single country are marked separately. A document is written in the language of the country where the problem is defined, so Korean problems exist in Korean and American ones in English, and we do not produce translations. Translation is the easiest way to smuggle in a judgement the original never made.

Research and writing are kept apart

When one party both researches and writes, the sources in the evidence table end up being self-reported: "I saw it" and "here it is" stop being distinguishable.

So only material the research step actually opened goes into a research brief, and the writing step may cite only what is in that brief.

We state exactly how far a machine enforces that. In an automated batch the brief and the body are compared, and a mismatch discards the entire day's batch — not just the one document. In a hand-run round that comparison is an editorial discipline rather than a machine check.

We also write down what this separation does not buy. There is no standing automated process that re-checks, at this moment, whether the sources in a brief still open. Links rot. We fix broken links when we find them, but nothing yet tells us first.

Machines refuse first

Two grades of check run when a document is loaded. Both run on every load, whatever route the document was written by.

An integrity violation in even one document aborts the entire load — not a single row is written. That covers a population estimate whose four figures disagree, a document without exactly thirteen parts, a missing required field, or a status line that cannot be read.

A document that trips a publication gate is the only one withheld from the site — attribution that does not stand up, a quotation over the permitted length, an image where there should be none, a computation file that cannot be parsed, traces of the authoring tool left in the body, and the like. The decision is fixed at load time, and the layer that serves pages does not reinterpret it.

About that last reason. In August 2026 an authoring tool's internal markup leaked into the end of a document, and what caught it was luck rather than a check. That class is now refused by name — separately from whether the content is right, we measure whether the content is the author's.

What we will not write

We do not use the real name of a private individual — whatever the source says. Once it is on the page that connection cannot be recalled. A placeholder such as "Official A" takes the place of the name.

Assertions of illegality, misconduct or discrimination appear only as citations of a public body's determination, with that body's name and the date of its decision in the same sentence. Repeating a claim made by an advocacy group or a press commentary is not citation; it becomes our assertion.

We describe the gap and do not supply a motive. We do not assert why a gap was permitted — capture, concealment, indifference. Thirteen parts set facts side by side, and the moment a motive is laid on top, an arrangement becomes an argument.

We do not name individual private companies in our own voice. The unit of this record is an institution, not a firm. A structural description takes the place of the name — "the two largest delivery platforms." Company names remain exactly where they belong: in the source column of the evidence table, and in the URLs.

We publish no images. Photographs and charts mostly belong to third parties, and not publishing them is more honest than adjudicating rights document by document.

We do not write what a minor should not see. This is not a rule about subject matter — child abuse, drugs, suicide, war and sexual violence are among the most important things this record covers. What is being separated is not the subject but the sentence.

Blanks carry meaning

When a part cannot be filled we record which of four reasons applies. Not collapsing them is the point — measured across the corpus, "we did not investigate this" ran three times as common as "the source does not say," and merging the two makes our own shortfall read as a failing of the sources.

Absent in source — the material we opened does not address it.

Not investigated — we could not establish the conditions for a judgement. This one is ours.

Not yet occurred — the point at which the item could exist has not arrived.

Not derivable — no chain could be built to size the affected population. Here we do not invent a number; we write down what was missing instead.

The "Missing N" on each card counts the parts marked this way. We do not reduce it by filling parts with self-evident sentences — there are documents where the filled version reads worse than the blank one.

What this record does not claim

Most documents are compiled from secondary material. Each one states its evidence type and authoring mode, and usually says "secondary" and "derived from press reports." We do not reword those to sound less derivative — that would be claiming an originality the sources do not support.

We do no field reporting. We do not interview people or verify conditions on site.

We do not declare any individual document solved or unsolved. Instead each says what would count as solved, so that rather than choosing the test ourselves we leave the reader to apply it. That is why the resolution field reads "not confirmed" on every document. It is not laziness; it is the limit of what this record can say.

Tools draft it; we state what a person is answerable for

Language models are used to produce drafts. We write that down rather than hiding it. For a period that drafting also ran automatically on a schedule.

So we state what a person actually carries, independently of whether that schedule is running.

Even while it runs automatically, the only thing it can change is document files — if a single byte of code changes, that day's output is discarded whole. The behaviour of this service, and the checks above, are out of automated authoring's reach.

Editorial judgement is exercised by a person, round by round — what to cover, how the rules should change, and whether something already published should come down.

Here is what that review actually refused. In August 2026 we re-researched 36 documents that carried no source URL at all. Thirty were rewritten around sources we found; six were taken down because no public evidence supporting their claim could be found. The corpus went from 256 documents to 250. A review that cannot decide in the subtracting direction is a review in name only.

How to get something corrected

If something here is factually wrong, tell us. Naming the document and the sentence, and saying why it is wrong, is the fastest route.

Once confirmed we correct it or take it down. Where we cannot find grounds to correct it, the document is taken down rather than repaired — that is what happened to the six above.

For rights infringement, the relevant clause of the terms of service sets out the procedure and the address that receives such notices.

Who publishes this

ML Labs · published by Kyengwhan Jee · ponderloft@gmail.com

The designated agent registration under US copyright law, including postal address and telephone number, is published in the "Designated agent under US copyright law" clause of the terms of service.

Related documents

AboutPrivacy policyTerms of Service