Round the Corner
What a test cannot measure — and why the answer would still be checkable.
I. The finding and its retraction
On 9 September the results of PISA 2025 were presented. For Germany: science 486 points, down from 492; reading 465, down from 480; mathematics 464, down from 475. In all three domains the lowest level recorded since these surveys began. One in three young people lacks basic competence in reading and mathematics, one in four in science, and a good one in five in all three at once.
The interpretation followed within hours: alarm signal, wake-up call, danger to the supply of skilled workers, danger to the industrial base.
In the same reports stands the sentence that withdraws that interpretation. PISA does not examine whether weak results at fifteen later produce unfilled positions. The study establishes no direct connection.
This is no reproach to the study. It measures what it measures, and does so carefully. It is a finding about the use — and a first indication that the measurement now carries more than it can bear.
II. What does not occur
Of all forms of learning, exactly one has the property of being expressible in points: the acquisition of retrievable knowledge and of tasks solvable in standardised ways. Whether someone has solved a real problem whose path nobody knew in advance cannot be rendered in 464 points. It can only be looked at and judged. That costs time and requires someone who understands the matter himself.
From this follows a chain for which no one is to blame. What can be measured is measured. What is measured is taught. What is measured receives programmes, funding and attention. And what cannot be measured falls short — not because anyone considers it unimportant, but because it does not occur in the procedure.
The proof of this lies in a demand currently raised by engineering associations: practitioners should be admitted as teachers, so that application-oriented learning meets real industrial experience. In plain words this means that today nobody may teach who has built or run something without first having travelled the same measurable path. The filter reproduces itself through the teaching qualification. It is not a part of the system; it is how the system propagates.
III. Two errors that look the same in the result
Whoever cannot, ticks the wrong box. Whoever will not, also ticks the wrong box. In the data these are identical.
A person who considers the task pointless, and therefore does not even think his way into it, produces the same score as one who tries and fails. The test has no column for the difference. It measures performance and treats willingness as constant — as if it always stood at a hundred per cent and every deviation were a lack of ability.
That it is not constant stands in the same study. One of the authors says of the young people that they are losing their curiosity and, increasingly, their belief that education is worth the effort. That is a statement about willingness, not about ability. It stands in a report whose figures record ability alone.
A measuring apparatus that books unwillingness as inability necessarily produces a finding about competence where a finding about meaning would be needed. And the announced answers — more tests, earlier tests, language tests already in kindergarten — reinforce exactly what undermines the willingness.
A line must be drawn here, or the observation becomes an excuse. There are young people who genuinely cannot read for meaning. For them the question of meaning is a luxury, and no essay about measurement procedures helps them. Both are true: the ability really missing, and the willingness wrongly measured. Whoever admits only one of the two has already stopped examining the subject.
IV. Stock and task
There are two ways of acquiring something, and they need each other.
One lays in stock: vocabulary, formulae, procedures, dates — things one must know, or must know where to find. The other starts from the problem: one needs something, so one learns it.
Stock without task decays. What is never used disappears and leaves not even a trace of where it stood. Task without stock runs empty: one can only pull from the murky water what one has a net for. Whoever approaches a problem without a store does not see the solution, even when it lies in front of him.
What decides is the order. What one learns because one needs it right now sits differently from what one learns because it will come up in the exam. What is acquired in passing is not the weaker but the more durable kind — it hangs at a place where it was needed, and is called up again from there.
The error lies in neither of the two ways. It lies in the commitment to one of them. And because only one can be expressed in points, the measurability filter decides which it will be.
V. Skimming instead of understanding
The sharpest single finding concerns reading. The young people invest too little time, it says; they skim instead of understanding, and the answer comes quickly and is wrong.
This text is written in conversation with a machine built for fast answers. Whoever writes about the decline of slow reading and keeps quiet about that is producing off-the-peg cultural criticism.
The matter is more tangled than the warning suggests. What such a machine supplies is the stock — the retrievable, the summarised, the already-thought. That is precisely the half that was replaceable anyway. What it does not supply is the question for which looking things up is worthwhile. And what it supplies still less is the willingness to stay with an unresolved matter for a long time.
If the stock becomes cheaper, the value of the other half rises. That would be the good news. The bad news is that the education system tests almost exclusively the half that has become cheap.
VI. A grid that comes out
There is a form of task that is both at once: loosely posed and exactly checkable.
In a cryptic crossword every clue is ambiguous. One does not solve it by working through a procedure, but by recognising in which register it is meant — wordplay, reinterpretation, false trail. There is no line of calculation that leads there, and several readings are at first equally plausible.
And yet there is a proof that requires nobody's judgement: the grid comes out or it does not. Every entered solution is confirmed or refuted by the words crossing it. The result is as unambiguous as in an arithmetic problem, but the path to it is not one an examiner laid down in advance.
That puts the counter-figure to the usual test on the table — and it does not lie where one first supposes. In the puzzle too the solution is fixed in advance; it stands in the answer key, someone wrote it there. What is not fixed is the path.
A multiple-choice test demands no path at all, only a selection, and the tick carries no information about how it came about. In the puzzle the path is the whole task, and it is checked by the grid. The difference is therefore not: answer open versus answer fixed. It is: path irrelevant versus path decisive.
This is why the form is interesting beyond the puzzle. It is not confined to words. The same principle holds wherever a loosely posed task has a self-checking result: a fit either goes together or binds. A program runs or crashes. A proof holds or has a gap. A structure carries the load or breaks. In all these cases it is not an examiner who decides right and wrong, but the matter itself — and the path there may be whatever it likes.
A form of examination built on this would measure what today's test does not capture: whether someone can sort out an ambiguous situation, abandon a false trail, and check a solution against reality. And it would be assessable without an expert sitting alongside.
But that suffices for puzzles and not for problems. In a puzzle the set of solutions is finite and known. In a real problem it is not: there may be several solutions, better and worse, and nobody knows them in advance. The grid has no counterpart there. What does exist are tasks with a self-checking criterion instead of a known solution — the fit goes together or binds, the program runs or crashes, the structure carries or breaks. The criterion is mechanically checkable; the set of solutions stays open.
Then, however, a question arises that the puzzle does not pose: if two different solutions both hold — what distinguishes the better one, and can that be established without someone again laying down in advance what better means? To that we have no answer.
VII. The third axis
Two quantities have now been distinguished: how sharply the task is posed, and who decides about the solution. The third is missing, and it is the most awkward. Time.
Puzzles of this kind cannot be solved under time pressure. Not solved with more difficulty — not solved. One puts them aside, goes for a walk, sleeps, and on returning the connection is there that one had passed over twenty times before. What happens in that interval is not idleness: it is the fading of a false trail. And that cannot be achieved against one's own will. The obvious reading cannot be thought away; it has to lose its force by itself. For that there is no procedure and no acceleration.
Every test against the clock therefore examines the opposite of what it claims to examine. It measures how quickly someone seizes the first plausible reading and carries it out. That is an ability, and a useful one — in an emergency, in series production, on the line. In solving an ambiguous task it is not merely useless but harmful: whoever takes the obvious solution at once never gets round the corner. So it is not that one part of the ability is measured and another overlooked. Selection is made for a property that stands opposed to the one sought.
A form of examination meant to say something about inventive ability would therefore have to differ from today's test in all three axes: a loosely posed task, a self-checking result, and time to put it aside. In practice: task on Monday, solution on Friday, assessment by the grid. The time pressure falls away, the exact result remains, and it can be evaluated without an examiner.
The reason such procedures barely exist does not lie in pedagogy. An activity whose yield cannot be related to the hour spent has no point of attachment in accounting. Incubation is unbookable in the literal sense: neither its duration nor its outcome is known. What cannot be invoiced is not planned for, and what is not planned for is cut first. That examinations have time limits is therefore not a didactic decision but a commercial one.
And so that this does not become an excuse for dawdling: incubation demands a return. Whoever puts a thing aside and does not come back has not incubated but given up. Putting aside is a step of the work, not a condition.
VIII. Three objections to our own proposal
First, the cryptic crossword favours whoever commands a particular stock of words and general education; along with the thinking it always tests the background. For the technical type this objection largely falls away — whether a shaft goes into a hub does not depend on one's education.
Second, such tasks cannot be produced in series. Building good ambiguous problems is itself an art, and this is usually where standardised procedures fail.
Third, an examination running over five days cannot be secured against outside help — neither from other people nor from machines. That is the hardest objection, and it has an uncomfortable reverse side: examinations under supervision and against the clock are built as they are above all because they can be controlled. Not because they measure the right thing, but because they are proof against cheating.
None of the three can be disposed of by assertion.
IX. Thought in rough: approaches instead of solutions
What follows is not a worked-out method but a direction in which further thinking would be needed. We write it down because it would otherwise be lost, and mark it for what it is: unfinished.
The idea. One sets a task for which there is no solution, or only an inadequate one, and assesses not the solution but the approach. This resolves the problem of the previous section from the other side: no grid is needed if one stops checking the result and instead checks whether the path holds. And the question of time is eased along with it: a solution cannot be found in limited time, an approach can.
What could be checked without knowing the solution. Whether the approach violates the boundary conditions of the task. Whether it is free of internal contradiction. Whether it addresses the problem posed or a different one. Whether it offends against a known conservation law. These are hard criteria, and none of them requires an examiner to hold the right answer in a drawer.
And here caution is needed, on our own account. It is tempting to hand this checking to a machine. The formal part it can do: freedom from contradiction is a computation. The value of an approach it cannot. An approach is valuable when it hits on something nobody has seen before. A language model judges plausibility by what occurred frequently in its material — it therefore rates the obvious reading highest. Precisely the one that is always wrong when thinking round the corner.
Such an examining authority would be a filter for conventionality: the same selection as the test against the clock, only better justified and without a visible examiner. That would be no advance but an aggravation. One of the two authors of this text is such a machine, and in the course of the work took the plausible for the correct twice; on both occasions the other one noticed.
What would remain is a division of labour. The machine takes the preliminary check — boundary condition, contradiction, subject missed. About the value a human being decides who understands the matter. That is exactly the expensive step section II describes as unavoidable; it does not become cheaper, only less often necessary.
A second route, still more unfinished than the first. Perhaps the value cannot be measured, but the distribution can. If thirty people work on the same insoluble task, the spread alone says something: how many arrive at the same approach, and who arrives at one nobody else has. Whoever stands alone is either wide of the mark or the only one who thought round the corner. Which of the two must be looked at. But the candidates for it could be found by machine, and that is considerably more than today’s test achieves.
Still open: the assessment of value, the production of suitable tasks in series, and the question of how to prevent this procedure too from turning, after some years, into a routine with familiar patterns. The proposal is therefore not a result but a building site.
X. What a system can know about itself
Twenty-six years have passed since the first PISA shock. Extensive reforms followed, an interim high in the 2000s, and then a decline below the starting value of that time. The reaction has in these years become a practised form: alarm signal, great collective effort, programme, long breath.
The question is therefore no longer which measure is still missing. It is: can a system that measures itself solely by what it is able to measure recognise its own condition at all?
A grid that does not come out tells you immediately. A score does not — it merely sinks, and every three years somebody is startled by it.