Rating dimension
Originality AnalysisWhat in a text is demonstrably its own work?
The check shows, at specific passages, which wording already appeared elsewhere and which factual claims were already published when the text ran. Every deduction points at a source you can open yourself.
“Load text” fetches the page's body text; you can edit it afterwards. The address also excludes the text's own page from the comparison sources — the check runs without one.
Check a text
Body text only, without navigation or captions. At least 400 characters, at most 60000.
The check still runs without a date. It just leaves no way to establish which source is the older one, and the report is then issued explicitly without an assessment.
Measured:
- how much of the body text matches older, retrievable sources word for word without naming them
- which checkable factual claims were already demonstrably published before the text appeared
- how many resolvable research artefacts the text carries — named primary sources, its own on-record interlocutors
Not measured:
- whether a text is good, true, important or relevant
- whether a text was written with the help of AI
- whether a text sits in the data a language model was trained on — that is not verifiable, so we do not assert it
- whether any copyright claim exists — that is a legal assessment, not a measurement
Three things that stay separate
A text can be entirely its own in wording and entirely familiar in substance. Averaging those into one figure hides exactly the case that matters. So the axes stay separate and are reported individually.
Wording
Do passages match older sources word for word without naming them? Evidenced by passage, source and word count.
Caps the result, never raises it
Factual claims
Were the text's checkable claims already published? Checked against dated sources, with the supporting passage.
Enters the assessment
Own research
Do the claims rest on named primary sources and on-record interlocutors that appear in no comparison source?
Enters the assessment
Topic density
How many newsrooms covered the same story in the same window? Reported as a plain count with a source list.
Deliberately does NOT enter the assessment
A story forty outlets covered does not make any single piece worse. Weighting it would penalise the routine coverage no newsroom can avoid.
How the check works
- 1
Search phrases are computed, not chosen
The rarest word sequences are identified and spread across the text by a fixed rule, not by a language model's judgement. The same input always yields the same queries, and every query issued is reproduced verbatim in the report.
- 2
Web search nominates candidates
The phrases are issued as exact-phrase searches. What comes back is a list of possible comparison sources — nothing more. A search hit is not yet evidence.
- 3
Every source is fetched and archived
Candidates are retrieved and stored with a retrieval timestamp and a checksum, so the finding stays verifiable even if the source page is edited later.
- 4
The comparison happens here, not in the search engine
Overlaps are computed locally between your text and the archived document. Only what survives that step appears in the report — with passage, source and match length.
The most important sentence on this page
Finding nothing is not proof of originality. Sources behind paywalls, printed originals, broadcast pieces without a text version and very recent publications are invisible to a web search.
So a text for which nothing could be retrieved does not receive a good assessment — it receives none, with a stated reason. A system that reads absence as a top mark systematically rewards topics that are hard to find.
Every figure carries its evidence tier
Evidenced
A retrievable source and a passage you can read for yourself.
Computed
A measure we calculate, with a stated failure mode — reproducible, but not a document.
Machine judgement
An assessment by a language model, always with its stated reasoning. It can move a value only inside its evidence class, never across the boundary.
That hierarchy is not just a display convention, it is built into the arithmetic: the entire correction budget of all softer signals combined is smaller than a single class step. A machine judgement therefore cannot justify a top mark, nor carry a downgrade.
Why there is no figure yet
The calculation is settled and described above. What is missing is the anchoring: thresholds and weights have to be calibrated against a reference corpus of German-language texts whose provenance is known. Until that exists the constants would be guesses — and a guess is indistinguishable from a measurement once it is on the screen.
So what ships first is the findings report: matches, checked claims, queries issued, evidence tiers. That is the part that already holds up.
How we word things
A finding is a statement about a text, not an accusation against a person. So we name what was measured instead of interpreting it:
| Instead of a verdict | we write what is measurable |
|---|---|
| a verdict on how a newsroom works | "Verified word-for-word match of 42 words with an older source, with no discernible attribution." |
| a statement about origin in AI systems | we make no such statement — it would not be verifiable |
| "unique" or "exclusive" | "Not found in the sources searched. Searched were: …" |
AI-generated analysis: Rating and reasoning are produced automatically and are not individually reviewed by an editor.