A whole product category now exists to explain a word to you without making you leave the page: reading assistants, AI answer panels, hover-definition browser extensions, the summary that sits above the search results. All of them are aimed at the same moment, and all of them describe it the same way.
You know the moment. You are reading. You hit a word, a name, an acronym, a clause you do not know. You open a new tab. Twenty seconds or twenty minutes later you come back, and the paragraph you were in has gone cold.
The industry agrees on the problem. What it has converged on is a family of solutions that all make the same move: bring the answer closer. Put it in a panel beside the page. Put it behind a keystroke that floats an overlay above whatever you were doing. Put it directly into the search results so there is no second page to load at all.
Each of these is an improvement on the tab. None of them is what it claims to be, which is an answer that arrives without moving you. The panel is also a tab. It is a narrower tab, docked, with a faster animation — and it still takes your eyes, your cursor and your place in the text.
This paper is about the difference between closer and in place, why that difference is the whole product, and where the in-place answer is the wrong design and the panel is right.
Every answer surface solves the tab problem by getting smaller. The exit path gets longer.
What we already argued
Two earlier papers in this series set up the ground. The Web Was Built for Navigation, Not Comprehension made the structural claim: the web's primitives — the link, the page, the tab — are all instruments of going somewhere else, and comprehension is the one task that going somewhere else actively damages. Your Knowledge Gaps Are the Most Intimate Data You Have made the consequence explicit: the record of what you did not understand is the most revealing thing a reading tool could hold, which is a reason to be careful about where that record lives.
This paper is the design argument that follows from the first, and it is narrower than either. It is about a single mechanism: where the answer appears, and what that costs.
Three ways the answer got closer
The current generation of answer surfaces comes in three shapes. The labels and the specific interfaces change often enough that pinning this to a version number would date the paper within a quarter, so what follows describes the patterns rather than the products.
The docked panel. A column opens beside what you were reading and holds a conversation. Browser assistants do this; search engines increasingly do it too, inviting you to continue a query with an assistant in a surface that opens alongside the results. The reading area is not replaced — it is squeezed.
The launcher overlay. A keystroke floats a box above everything. This is the Spotlight lineage: a modal surface, summoned from anywhere, that expects a typed command and then disappears. Desktop assistants have widely adopted it because it is the fastest thing to build that is available system-wide.
The rewritten result. The answer is composed directly into the page you would have navigated to anyway — the search results page becomes the answer rather than an index of places the answer might be.
The third of these is genuinely different and genuinely good, and it is worth saying so plainly. It removes a navigation. But it only helps when the question was a search. It does nothing for the case this paper is about, which is a question that arose inside a document you are already committed to reading.
What the first two have in common
Both the panel and the overlay solve the tab and leave the cost.
Consider what actually happens when you use them. Your eyes leave the sentence. Your attention re-anchors on a new region of the screen. You read an answer composed without reference to the specific sentence that confused you, because you had to restate the question in a box. Then you go back — and going back is the expensive part. You have to find the sentence again, reload the argument up to that point, and recover the reason you stopped.
The tab was never the cost. Re-entry was the cost. The tab was just the most visible carrier of it.
A panel reduces the distance you travel and leaves the re-entry intact. An overlay is worse on one axis: it covers the thing you were reading, so the context you needed to formulate the question is occluded at the exact moment you are formulating it. This is why the launcher pattern, which is excellent for launching, is a poor fit for understanding. Spotlight is optimised for "go do a thing." It has no notion of the thing in front of you. Asking it to help you read is asking a doorway to be a desk.
None of this is a claim about how many seconds or minutes anything costs. There
is a real research literature on task switching and on resuming interrupted work,
and it is the right literature to settle the size of this effect. Neither the
author nor this paper holds a defensible number for the specific case of an
in-document lookup, and the widely circulated figures on "time to refocus" are
routinely quoted far outside the conditions they were measured in. The
directional claim — that resumption has a cost and that cost is separate from
travel time — is the one this paper makes. CITATION-NEEDED: task-resumption literature, specifically any study measuring resumption lag for a lookup initiated from within a reading task rather than an external interruption.
What "in place" would actually require
If the mechanism is re-entry, then the design target is not proximity. It is not having left. That is a stricter specification than it sounds, and most things that call themselves inline do not meet it. Four conditions:
1. The context is taken, not retyped. The question is raised by a selection. The tool already knows the sentence, the surrounding paragraph, and the document they sit in. If the user has to restate what they are asking about, the tool has already made them leave — restating requires holding the context in your head instead of on the screen.
2. The answer is anchored to what raised it. It appears at the selection, not in a region reserved for answers. The distinction matters because an answer in a fixed region has to be found; an answer at the selection is already where your eyes are.
3. Nothing is displaced. The paragraph stays put. A layout that reflows to make room has moved you, because the sentence you were on is now somewhere else. This rules out both the docked panel that squeezes the column and the overlay that covers it.
4. Dismissal leaves no residue. When the answer goes away, the page is exactly as it was, and your eye is on the word you stopped at.
Meet all four and the interaction stops being a lookup and becomes something more like an annotation you happened to summon. Miss any one of them and you have built a faster tab.
The second question is the tell
Here is the part that changed my view of what this product is, and it came from using the thing rather than designing it.
An explainer that answers once is a dictionary. The first answer is almost never the end of the matter — it resolves the word and raises the next thing. What you actually want to say is "yes, but why does that apply here", and v1 could not take that question.
Now notice what the surfaces above do to follow-ups. In a panel or an overlay, asking again is cheap within the panel and expensive to relate back to the text. So people compensate by asking one large, well-formed question — because the switch has already been paid for, you may as well make it count. In place, the incentive inverts. Asking is so cheap that you ask three small questions instead of one big one, and the small ones are better: they are specific to the sentence, they build on each other, and each one is answerable.
The cost of asking determines the shape of what gets asked. That is the real consequence of where the answer appears, and it is not a matter of seconds saved. A tool that makes follow-up free produces a different kind of understanding than one that makes it merely possible.
Where in place is the wrong answer
The claim above is bounded, and the boundary is not a footnote — it is most of the map.
In-place answering is for comprehension in flow: you are committed to a document, and something in it is blocking you. That is a real and frequent situation, and it is badly served today.
It is the wrong design for everything else. Comparing four sources, holding a table open while you read against it, building up a synthesis over an hour, anything where the answer is the artifact you are working on rather than an unblocking — for all of these the panel is correct, because you want the answer to persist, to have room, and to be somewhere you can return to deliberately. A popover that politely vanishes is exactly wrong for work you intend to keep.
So this is not an argument that the panel is a mistake. It is an argument that the panel has been asked to cover a case it structurally cannot, and that the case it cannot cover is the common one.
What would settle it
The honest position is that this paper argues from mechanism and from use, not from measurement. The study that would settle it is not difficult to specify: same readers, same documents, same set of genuinely unfamiliar terms, and a comparison of comprehension and completion between an in-place answer and a docked panel — with follow-up question count recorded, since the prediction here is specifically that in-place raises the number of questions asked and lowers the size of each.
The prediction is falsifiable and I have not run it. Stating that plainly is worth more than a number that would not survive being checked.
This paper argues the mechanism. The companion, The Second Question, is the design account: what months of using an in-place explainer surfaced, why the form factor had to leave the browser, and what a standalone version has to be.