← Writing · Civic & Democratic Infrastructure
Flux Working Paper No. 37

Who Gets to Say It's There

Ken Ruto · Flux (FluxImpact) · September 2026 · 24 min · Updated Sep 2026
Revision history
2026-09-18 — full white-paper revision: scope/standing section, four-way decomposition of disagreement, the claim-record mechanism, objections, falsification conditions, evidence table and citation register.
Read as paper ↗
BibTeX · RIS
Civic & Democratic InfrastructureOffline-first / every fact sourced

The first paper in this series promised a model of the city "residents can change, not just look at." The most recent one is titled Legible, Navigable, Writable and spends its length on measurement: heights from satellite imagery, and the discovery that two datasets which both call themselves measured disagree by nearly seven metres.

That paper delivers legible. It is good work and I stand behind it. But the third word in its own title has now been promised twice and delivered zero times, and a series that has been this careful about admitting what it does not know should not end on an unexamined adjective.

So: what does writable actually require? Having spent a while on the question, I think the honest answer is that the write itself is trivial, and everything that makes a writable city hard begins with the second write.

The claim I want to defend is narrower and more uncomfortable than "conflict resolution is hard." It is that the three available ways of resolving a conflict between two claims about a city are not three implementation options. They are three different theories of who owns the record, each of which imports an institution, and in a city where the official record is thin and unevenly distributed only one of them can be shipped without first building an institution that does not exist. That is not an engineering conclusion. It is a governance one, and this series has been using a governance word as though it were a feature flag.

KANAIRO://WP37 — TWO WRITES, ONE BUILDING LEGIBLE IS A DATA PROBLEM. WRITABLE IS A GOVERNANCE PROBLEM. INCOMING EDITS RESIDENT · 4 FLOORS SURVEYOR · 11 M LANDLORD · 6 FLOORS SATELLITE · 17 M COUNTY · NO RECORD BUILDING 4417 FOOTPRINTAGREED HEIGHT11 M / 17 M CONTESTED USECONTESTED OWNERCONTESTED ARBITERNONE APPOINTED THE WRITE IS THE EASY PART. THE SECOND WRITE IS THE WHOLE PROBLEM. WP27 PROMISED A CITY RESIDENTS COULD CHANGE. THIS IS WHAT THAT COSTS. The write is the easy part. The second write is the whole problem.

1. Scope, method, and standing

This paper reports no built write path. There is no way for a member of the public to change a building attribute in the Nairobi twin today, beyond the billboards discussed in §2, and nothing in what follows has been tested against real contributors.

What it does is analytical, and rests on three things. The first is the series' own record: six papers of measurement, coverage and estimation work, whose findings I take as established within this corpus and cite by cross-reference. The second is the published literature on volunteered geographic information, where the questions this paper asks have been asked for two decades and answered with evidence I do not have — I lean on it where it is load-bearing and cite it. The third is an argument about institutions, which is mine and is the part to attack.

I should also state the standing problem, because it shapes what this paper is allowed to conclude. I am not a resident of the settlements whose representation is at issue in §5. A paper arguing about who should have authority over a record of somebody's home, written by someone who does not live there and has not asked them, can make an argument about mechanism but cannot settle a question about legitimacy. §12 keeps to that line. Where the paper reaches a conclusion that a resident consultation could overturn, I say so.

2. We have already shipped a write primitive, and did not notice

It is worth noticing that this series began with one and did not recognise it as such. The $1.99 billboards in the first twin paper — sold as a joke — are a write: a member of the public places a persistent object into the shared model and everyone else sees it. The joke worked precisely because the mechanism worked.

What made the billboards tractable is that they are non-rival and unfalsifiable. Your billboard does not contradict mine. There is no fact of the matter about which is correct, so there is nothing to adjudicate, and the model can accept every write in the order it arrives.

Unpack that into the three properties actually doing the work, because they are the test any candidate write has to pass:

No shared referent. Two billboards are two objects. Two height claims are one object with two values. The moment two writes point at the same thing, the model has to decide what the thing is.

No truth condition. A billboard cannot be wrong. It can be ugly or misspelled, but there is no state of the world it fails to correspond to. A height can be wrong, which means somebody can be shown to be wrong, which means somebody has to do the showing.

No stake. Nobody's rent, tenure, rates assessment or planning permission turns on what a novelty billboard says. This is the property that is easiest to forget and hardest to engineer around, and §8 is about what happens when it fails.

Almost nothing else about a city has all three. Most of what would make the twin useful has none of them.

3. The second write

Consider the single most useful thing a resident could tell the model: how tall this building is. We already know that for 99.7% of Nairobi's 1.27 million buildings nobody has recorded a height, and we filled that silence with a confidence-tiered estimate rather than a confident guess.

Now let a resident write the real number. The first write is a straightforward database operation. Then the landlord writes a different number. Then a surveyor writes a third, and the satellite-derived estimate — which the measurement paper showed can itself be off by metres — says something else again.

The model now holds four claims about one building and no procedure for choosing among them.

It is worth being exact about why these four disagree, because "data quality" flattens four different situations into one and each has a different remedy:

They measured different things. Height to eaves, height to ridge, height above ground on the uphill side, storeys times an assumed storey height. In a city built on a slope this is not pedantry; it is most of the seven metres.

They measured at different times. A building that gained a floor between two observations produces two correct numbers. Nairobi's building stock is not static, and a model without a time dimension turns change into disagreement.

One of them is wrong. Ordinary error, in the instrument or the transcription.

One of them is lying. Rare, consequential, and the case every system claims it will handle later. §8.

Only the third is a data-quality problem in the usual sense. The first is a schema problem the series can fix and has not. The second is a modelling problem. The fourth is a governance problem, and it is the one that determines the architecture, because a design that handles the first three and fails on the fourth is a design that works right up until the model starts to matter.

That is not a storage problem or a UI problem. It is the question of who has authority over the record, and it is the reason a writable city is a different kind of object from a legible one.

4. Three ways to resolve it, and what each costs

Authority. The county decides; residents submit, the county ratifies. This is how land records work, and it has the advantage of matching an existing institution. It also reintroduces exactly the bottleneck the twin was supposed to route around: the settlements missing from the map are missing in large part because the official process never reached them, and an official ratification step puts that process back in the critical path.

There is a second cost that is less obvious and, I think, larger. Authority does not merely slow the record down; it changes what the record means. A ratified height is an assertion by the county, and assertions by the county have consequences attached — for rates, for planning enforcement, for tenure. A resident deciding whether to contribute a measurement is then deciding whether to invite those consequences. The mechanism that makes authority trustworthy is the same mechanism that makes participation risky, and it will be most risky exactly where the record is thinnest.

Consensus. OpenStreetMap's model: contributors argue, the last edit standing wins, disputes escalate to community norms. It genuinely works at scale, and it is the reason we have any footprint coverage at all.

But its outcomes track contributor density, which is precisely the bias the third paper in this series measured — the difference between two settlements' coverage had nothing to do with the settlements. This is not a local finding. Herfort and colleagues, analysing OSM building data across more than thirteen thousand urban centres, found completeness above 80% for around 16% of the world's urban population while remaining below 20% for cities holding roughly 48% of it, with the gap structured by development index, city size and region rather than by anything about the cities themselves.1 Humanitarian mapping has narrowed the gap and not closed it.

The consequence for a conflict-resolution rule is the part usually left implicit. Consensus among an unrepresentative set of contributors does not produce a neutral answer with a known error bar. It produces an answer whose error is correlated with exactly the thing you are trying to measure, and then presents it without a marker. That is worse than a wrong number. It is a wrong number that has been laundered into a fact.

Provenance. Keep every claim, attach who said it and how, resolve nothing. The model stops returning "the height" and starts returning "four claims, here is each one's source and confidence."

This is not a novel invention and I do not want to present it as one. The W3C's provenance model — entity, activity, agent: what exists, what produced it, who was responsible — has been a Recommendation since 2013 and is the standard vocabulary for saying exactly this kind of thing about data.2 The contribution this paper can make is not the data model. It is the argument in §5 about why, for this city, it is the only one of the three that can be shipped.

KANAIRO://WP37 — WHO DECIDES, AND WHAT IT COSTS THREE WAYS TO SETTLE A CONTESTED WRITE AUTHORITY WHO DECIDES THE COUNTY RATIFIES WHAT IT COSTS PUTS THE OFFICIAL PROCESS BACK IN THE CRITICAL PATH THE SETTLEMENTS MISSING FROM THE MAP ARE MISSING BECAUSE OF THAT PATH CONSENSUS WHO DECIDES CONTRIBUTORS ARGUE WHAT IT COSTS OUTCOMES TRACK CONTRIBUTOR DENSITY, NOT GROUND TRUTH LAUNDERS THE COVERAGE BIAS THIS SERIES ALREADY MEASURED INTO A FACT PROVENANCE WHO DECIDES NOBODY — KEEP ALL WHAT IT COSTS THE MODEL STOPS RETURNING A NUMBER. EVERY CONSUMER DECIDES NEEDS NO INSTITUTION NAIROBI DOES NOT HAVE. THE ONLY SHIPPABLE ONE. THE HEIGHT HEURISTIC ALREADY CHOSE THE THIRD COLUMN. THIS EXTENDS IT FROM ESTIMATES TO WRITES. RESOLVING NOTHING IS A POSITION, NOT AN ABSENCE OF ONE. Authority needs an institution. Consensus inherits a bias this series already measured. Provenance resolves nothing, on purpose.

5. Provenance is the only one we can actually ship

Not because it is elegant — it is the least satisfying of the three — but because it is the only one that does not require an institution we do not have.

Set the three side by side against what each presupposes:

Regime What it presupposes Present in Nairobi?
Authority A cadastral institution with reach into every settlement, and residents who can contribute without inviting enforcement Reach is the thing the series measured as absent
Consensus A contributor population whose density is not correlated with the attribute being measured Documented to be correlated, globally and here
Provenance An identity for each claim good enough to attribute it, and consumers willing to handle contested values Identity is hard but tractable; consumer burden is real and is the cost

Authority and consensus both require something to already be true about the city before the mechanism works. Provenance requires something to be true about the consumers of the model, which is a thing a builder can actually negotiate.

It is also the option this series has already committed to twice, which is the part I find persuasive. The height heuristic paper attached a visible confidence label to every estimate rather than shipping a flat city or a silent guess. The measurement paper's central finding is itself a provenance finding: two sources disagree, and the one everybody treats as authoritative is probably the worse of the two — a conclusion only reachable because both were kept and labelled rather than reconciled into one number.

Extending that from estimates to writes is a small step conceptually and a large one in what the model is for. A twin that returns contested values is harder to build applications on. You cannot ask it "how tall is this building" and get a number. You have to ask "what is claimed about this building, by whom" — and every consumer downstream has to decide what to do with that.

That is a real cost, and I do not want to wave it away. It is also, I think, the only version that is honest about a city where the official record is thin, contested and unevenly distributed. The alternative is to pick a winner and hide the disagreement, which is exactly the failure the measurement paper caught somebody else committing.

6. What this looks like as a mechanism

The position in §5 is only useful if it survives contact with a schema, so here is the concrete shape, stated plainly enough to be argued with.

The unit is a claim, not a value. A claim carries: the subject (which building), the predicate (which attribute), the value, the definition under which the value was produced (height to eaves, to ridge, storeys × assumed storey height), the observation time, the method, the origin, and a confidence. Seven fields where the current model has one. Three of them — definition, time, method — exist to retire the first two disagreement types in §3 before any conflict machinery runs, which is the cheapest win available and does not require a write path to be worth doing.

The origin is a role, not a name. This is the part §8 says is unsolved, but the shape is clear enough to state: the field records what kind of party made the claim and what standing that gives it — resident of the structure, holder of a registered interest, licensed surveyor, automated estimate — plus a stable identifier that lets a series of claims be recognised as coming from the same origin. What it must not require is that the identifier resolve to a person for anyone who reads the record.

The API stops returning scalars. Today a consumer asks for a height and receives a number. Under this model it receives a claim set, or it names a resolution policy and receives a value plus the policy that produced it and a flag if the set was contested. The second form is what almost every consumer will use, and building it first is how the proposal avoids §8's adoption objection. The important discipline is that the contested flag is not optional and cannot be suppressed by the caller, because a caller that can hide the contest will, and then §5's whole argument has been implemented into a system that behaves like consensus.

Nothing is deleted. A superseded claim is superseded, not removed. This is what makes the record auditable later and it is the property that costs nothing now and cannot be retrofitted.

Two consequences worth naming. The storage cost is real but small — claims about buildings are tiny records and 1.27 million of them is not a large table. The interface cost is the one that will actually decide this, and it is the thing I have not designed: a contested value has to be shown to an ordinary reader in a way that helps rather than alarms, and every screen in the twin is a place to get that wrong.

7. What provenance does not mean

Because the word is doing a lot of work, four disclaimers about what the position in §5 is not.

It is not "never resolve." A consumer that needs a single number will always be able to ask for one under a stated policy — highest confidence, most recent, prefer-surveyed. The difference is that the policy is the caller's, declared at the point of use, and the underlying claims survive it. Resolution becomes a view rather than a write.

It is not a reputation system. Scoring contributors and weighting by score reintroduces consensus with extra steps, and inherits its bias: a reputation built by contribution volume is a reputation that tracks contributor density.

It is not neutrality. Keeping every claim is itself a choice with winners. A landlord and a tenant do not stand in the same relation to a persistent, attributable claim about a building, and §8 is about which of them the design favours.

It is not free. The storage is trivial; the interface is not. Every surface that displays a contested value has to display the contest in a way an ordinary reader can act on, and doing that badly produces a model people stop trusting rather than one they trust appropriately. This is the largest unbudgeted cost in the proposal and I have not designed it.

8. The adversarial case, which I cannot solve

The hardest version is not a resident misremembering. It is a party with a material interest in the recorded number, writing at scale.

Consider who actually has an incentive to move a building height. A landlord facing a rates assessment keyed to floor area. A developer whose approved plans and built structure have diverged. A speculator establishing a paper trail for a structure that would not survive inspection. In each case the writer is better resourced than the residents whose claims would contradict theirs, and in each case the contest is not symmetric: one side is writing about an asset, the other about their home.

Provenance makes this visible rather than preventing it. That is a genuine property and it is weaker than it sounds. Visibility works when somebody is looking and can act; a contested-value marker on a record nobody audits is documentation of a capture, not a defence against one.

Two partial mitigations I can see, neither tested:

Asymmetric cost of assertion. Claims that carry a stake — those made by a party with a registered interest in the parcel — could be required to carry stronger evidence than claims that do not. This is ordinary evidentiary practice and it is implementable. It also requires knowing who has a registered interest, which is the cadastral reach §5 says we do not have.

Attribution without identification. A resident should be able to make an attributable claim without being personally identified to a landlord — the claim needs a stable, accountable origin, not a name. This is the single most important unsolved design question in the proposal, and getting it wrong makes the tenant's position worse than saying nothing.

I raise both to be explicit that §5's conclusion is conditional on them. A provenance model deployed without an answer to the adversarial case is not neutral between the parties. It is a mechanism that is cheap for the well-resourced to use and risky for everyone else, which is the same shape of failure this corpus has documented elsewhere.

9. Objections I take seriously

"OSM already solved this and you are reinventing it badly." OSM has the richest edit history and the strongest conflict norms of any open geographic dataset, and I am not proposing to replace it. The objection I am making is not to OSM's machinery but to its outcome rule: a single agreed value per object, arrived at by a contributor population whose distribution is documented to be uneven.1 The changeset history preserves who said what and is the closest existing thing to what §5 asks for — it is just not what the API returns or what downstream consumers read.

"Contested values will kill adoption." Probably the strongest objection, and I have no counter-evidence. Every consumer wants a number. The honest response is that §7's first point exists precisely for this — a default resolution view with the contest one layer down — and that if even that proves unusable then the position in §5 fails on practical grounds, which §11 records as a falsification condition.

"This is a lot of theory for a model with no users writing to it." Fair. The defence is in §10: the choice between the three regimes is a data-model choice, and data-model choices are cheap before there are writes and expensive after. The theory is early because the decision is early.

"Why not just let the county do it and accept the exclusion?" Because the exclusion is not a side effect, it is the thing the series exists to document. The third paper measured who the map leaves out; a design whose first move is to route authority through the institution that produced the omission is a design that has decided the omission is acceptable.

10. What this changes about the twin

Four things, if the argument holds.

The record stops being a row and becomes a set of claims with sources — a change to the data model, not the interface, and one that is far cheaper to make before there are writes than after.

"Writable" stops meaning "has an edit button" and starts meaning "has a resolution policy." The edit button is a weekend. The policy is the product.

The question the series should be asking about a resident's contribution changes from is this correct to is this attributable. The first is often unanswerable in a city like this one. The second is almost always answerable, and it turns out to be enough to be useful.

And the schema acquires a time dimension and a measurement-definition field before it acquires a write path, because §3 says two of the four disagreement types are artefacts of their absence. Fixing those first means the conflict machinery only has to handle real conflicts.

11. What would falsify this

A city-scale twin ships multi-claim provenance and its consumers cannot use it. If downstream applications route around the contest by silently taking the first value, provenance has produced the appearance of honesty and the substance of an arbitrary choice, and §5 is wrong on practical grounds.

Consensus produces unbiased coverage somewhere the contributor population is uneven. §4's objection to consensus is empirical, and a counterexample would retire it.

Attribution without identification proves impossible at acceptable cost. Then §8's second mitigation fails, and with it the claim that provenance is safe to deploy in a city with a landlord–tenant power asymmetry. I would rather withdraw the recommendation than ship it without this.

The county's reach turns out to be adequate. If a cadastral process can in fact be run into the settlements at reasonable cost, authority beats provenance on every axis and the argument is moot.

12. Limitations, and what this paper does not claim

It does not report a built system. There is no write path in the Nairobi twin today beyond the billboards, and nothing here has been tested against real contributors.

It does not claim provenance is correct in general — for a city with a functioning cadastre, authority is obviously better, and the argument above is specific to a record that is thin and unevenly distributed.

It does not claim the three regimes are exhaustive. They are the three I can name and defend. Hybrid designs exist in other domains and I have not surveyed them.

It does not resolve the adversarial case (§8). Provenance makes it visible rather than preventing it. Whether visibility is sufficient is an empirical question this series has not earned an answer to yet.

And it does not settle the legitimacy question flagged in §1. Whether residents of a settlement want a persistent, attributable record of their buildings is not something this paper can answer on their behalf, and a recommendation that proceeded as though it could would be making the same error the paper accuses authority of making.

What it does claim is that the word was load-bearing and unexamined. Writable was never a feature we had not got round to. It is a governance commitment, and the series should either make it or stop using the word.


Appendix A — Evidence table

S peer-reviewed or standards source, C established within this corpus by an earlier paper, not sourced.

# Claim used in this paper Value Source
1 Nairobi buildings with no recorded height 99.7% of ~1.27m C, WP30
2 Disagreement between two "measured" height datasets ~7 m C, WP32
3 OSM building completeness >80% ~16% of world urban population S1
4 OSM building completeness <20% ~48% of world urban population S1
5 OSM coverage bias structured by HDI, city size, region qualitative finding S1
6 Provenance vocabulary: entity / activity / agent W3C Recommendation, 30 Apr 2013 S2
7 Coverage difference between two Nairobi settlements unexplained by the settlements qualitative finding C, WP29
8 Share of the ~7 m disagreement attributable to datum/definition rather than error
9 Rate of building-stock change in Nairobi (floors added per year)
10 Any city-scale twin implementing multi-claim provenance for building attributes
11 Willingness of settlement residents to contribute attributable building claims
12 Incidence of interested-party editing in an open geographic dataset

Row 11 is the one that should be collected first, and not by me — it is the question §1 says this paper has no standing to answer. Row 8 is the cheapest: it is a re-analysis of data the series already holds, and it would tell us how much of the conflict problem is really a schema problem.

Appendix B — Open questions and citation register

CITATION-NEEDED — prior art in multi-claim civic records. Whether any city-scale digital twin has implemented multi-claim provenance for building attributes, and what its consumers did with contested values. I could not find one while writing this, but absence of a find is not evidence of absence, and this should be checked properly before the claim that it is novel is made anywhere. Appendix A row 10.

CITATION-NEEDED — interested-party editing. §8 asserts that parties with a material stake would write to the record. That is a prediction from incentives, not a measurement. The literature on vandalism and undisclosed paid editing in open geographic data may already bound it; I have not read it.

CITATION-NEEDED — attribution without identification. §8's second mitigation assumes a claim can carry accountable origin without exposing a contributor to a counterparty. Anonymous-credential and verifiable-claim work in other domains likely bears on this directly, and should be reviewed before the design is attempted.

Internal re-check. Rows 1, 2 and 7 are this corpus citing itself. Before any of them is quoted outside Flux, each should be traced back to the underlying dataset and computation rather than to the paper that reported it.

  1. Benjamin Herfort, Sven Lautenbach, João Porto de Albuquerque, Jennings Anderson and Alexander Zipf, "A spatio-temporal analysis investigating completeness and inequalities of global urban building data in OpenStreetMap," Nature Communications 14, article 3985 (2023), doi:10.1038/s41467-023-39698-6.

  2. PROV-O: The PROV Ontology, W3C Recommendation, 30 April 2013, W3C Provenance Working Group. The model's three core classes — entity, activity, agent — are the standard vocabulary for recording what a claim is, what produced it and who is responsible for it.

Ken Ruto
About the author
Ken Ruto

Founder of Flux. Building vertical AI-powered SaaS for Africa's institutions — and writing the thesis behind every bet. kenruto.fluximpact.org →

Share X LinkedIn WhatsApp
This paper is about this

KaNairo is the live twin this series describes — 1.27M real building footprints, real streets, real traffic. Walk it yourself rather than take the paper's word for it.

Did this land?
Was it useful?

Comments

No comments yet — be the first.

Replying to · cancel
Get new essays

No spam — just the next piece when it's out.

Think I got something wrong? Highlight any sentence to push back on it — or It comes straight to me, never shown publicly.

Push back
The Nairobi twin series
22 min
The Twin Nairobi Doesn't Have Yet
We let people pay $1.99 for a joke billboard in a toy 3D Nairobi. It's a working model of the real feature: a twin residents can act on, not just admire — and what "1:1" actually requires.
7 min
From the CBD to All of Nairobi: A Tile-Streamed Building Pipeline for a City-Scale Twin
The v1 render covered 1,026 buildings in the CBD. This is how it became 1,277,511 real buildings across greater Nairobi, streamed tile by tile instead of loaded whole.
6 min
Who's Missing From the Map: Building-Footprint Coverage Gaps in Nairobi's Informal Settlements
Kibera: 82% of its buildings are in OpenStreetMap. Mathare: only 16% are. The difference is not population or need — it is whether a settlement ever had a dedicated volunteer mapping campaign.