← Writing · AI Automation
Flux Working Paper No. 39

The Switching Cost Is Why Nothing Sticks

Ken Ruto · Flux (FluxImpact) · September 2026 · 35 min
Read as paper ↗
BibTeX · RIS
AI AutomationOffline-first / every fact sourced

A friend of mine has, at last count, four wellness apps on her phone: one for meditation, one for water intake, one for a step count, one that wants her to log what she eats. She paid for two of them. She opened all four in their first week. She has opened none of them in the last month. She has not deleted them either — they sit there, icons slightly dimmed by disuse, a small museum of intentions.

This is not a story about willpower, and it is not a story about bad apps. The meditation app is well reviewed. The food-logging app was recommended by a doctor. What happened to my friend has happened to almost everyone who has ever tried to use a phone to become a better version of themselves, and the previous paper in this series noticed the pattern without quite naming what causes it: she already lives inside her messages — a family group, a set of DMs, a note-to-self thread she has kept running for years — and none of those four apps do.

That paper made a claim I still believe: the most useful personal AI will not be an app you remember to open, it will be a presence already inside the conversation you are already having. What it did not do is explain why that should be true as a matter of mechanism rather than taste. "People already live in their messages" is an observation about where attention sits. It is not, by itself, an argument about why moving a tool into that space would change what the tool actually gets used for. A skeptical reader is entitled to ask: maybe messaging just feels more comfortable, the way a worn pair of shoes feels more comfortable, and comfort is not the same thing as effectiveness. This paper is the argument that closes that gap, and it turns out the gap closes on a finding from cognitive psychology that has nothing to do with phones and predates the smartphone entirely.

KANAIRO://WP39 — THE DETOUR AND THE DIP STANDALONE APP INTENTION OPEN APP ACTIVATION DIPS DONE RESIDENT IN THE THREAD INTENTION DONE NO DOORWAY. NO DIP. THE DETOUR IS NOT THE DELAY. THE DIP IN THE MIDDLE IS. The detour is not the delay. The dip in the middle — the goal losing strength while you go find the app — is.

1. Scope, method, and standing

This is an argument paper built on three kinds of material, and I want to be exact about the weight each one is carrying before using any of it.

What it argues from. First, the published experimental literature on task interruption and resumption in cognitive psychology — a research programme running from the early 2000s to the present, with a specific, falsifiable model at its centre (§3). Second, published measurements of what actually happens to engagement once a personal app that requires deliberate opening is installed — mental-health-app usage panels, industry retention aggregates, and the chatbot-failure literature (§4, §6). Third, published usage data on messaging behaviour in Africa specifically, since this series writes Africa-first and the argument changes shape once you leave markets where "just download another app" is nearly free (§5). Fourth, an argument connecting the three, which is mine, is the newest part, and is the part a critical reader should attack first.

What it does not argue from. Rafiki, the Flux product this paper is ultimately about, carries the status in_research in Flux's own product record as of this writing. It has not shipped. Nothing in this paper reports usage data, retention curves, or user testimony from a deployment, because there is no deployment. Every claim about Rafiki specifically is a claim about what the argument in §3–§5 predicts it should do, not a claim about what it has done. Where this paper draws on Rafiki's product description, it quotes it once, exactly, and treats it as a statement of design intent rather than as evidence: "Rafiki is the personal operating layer: an AI that lives in your messages, not in an app you have to remember to open. It talks the way you talk to a close friend, acts on your behalf in the real world — booking, ordering, reminding, following up — and checks in on you proactively, the way someone genuinely in your corner does. Its job isn't to obey; it's to hold you to the person you said you wanted to be."1

On the central gap. I have not found, and did not expect to find, a single controlled study that takes one well-specified personal task — say, a daily check-in on a stated goal — and randomly assigns otherwise identical users to a standalone-app version and a messaging-native version of the same tool, then measures completion over months. That study, to my knowledge, does not exist. This paper's argument is therefore an inference from a mechanism (§3) plus a pattern (§4), not a report of a direct causal test. I register that gap formally in Appendix B rather than pretending the literature closes it, and §10 states what such a study finding either way would do to this paper's claim.

On press-reported figures. Where this paper cites platform-penetration statistics — WhatsApp's share of internet users in Kenya, Nigeria and South Africa; app-abandonment aggregates — these are industry-reported figures, not audited data, and are used to establish an order of magnitude and a direction, never a precise rate.

What would count as refuting it. Stated in §10, before the argument's conclusion rather than folded into it afterward, for the same reason this series always states it there: a falsification condition written last is a rhetorical flourish, not a commitment.

2. What "lives in your messages" actually claimed, and what it left open

Worth being precise about what the previous paper established, so this one does not spend its length re-arguing a point already made or, worse, quietly smuggling in a stronger version of it.

The Personal Operating Layer observed that people already use messaging threads as a kind of externalised working memory — texting themselves reminders, keeping a group chat with exactly one member, running a to-do list inside a DM because opening an actual to-do app never happened. It argued that the natural home for a personal AI is therefore not a fifth icon on the home screen but the thread itself, and that the assistant's job is less to be obeyed than to hold the user to commitments they made about their own life.

That is a claim about where the tool should live. It is not yet a claim about why living there works better than living somewhere else, and the paper mostly left that as intuitive. Intuitive is not the same as established, and the strongest version of the skeptic's reply goes like this: every failed productivity app in history was also, at the moment of installation, the thing the user chose because it felt right. Feeling like the natural place for a tool to live is exactly what installing a fourth wellness app also feels like, in week one. If "this is where my attention already is" were sufficient on its own, meditation apps — which live on the same home screen as messaging, one thumb-swipe away — would not have the retention problem they demonstrably have (§4). Proximity on the phone is not the same as residency in the conversation. Something more specific than "convenient" has to be doing the work, and it has to be specific enough to predict which tools fail and which do not, not just to explain the ones that already succeeded.

The rest of this paper is the attempt to name that something.

3. The switching cost, defined precisely

Cognitive psychology has a name for what happens when you form an intention, get interrupted, and have to come back to it, and the name is more exact than the everyday phrase "losing your train of thought" suggests.

The foundational account is Altmann and Trafton's memory-for-goals model.2 The model treats a goal — "log what I ate," "message the client back," "do today's check-in" — as a memory trace with an activation level, the same currency ordinary memories are stored in. Activation decays over time and is boosted by rehearsal and by environmental cues. When you switch away from a goal to do something else, its activation does not freeze in place waiting for you; it decays, exactly the way an unrehearsed fact decays. Coming back to the goal later means retrieving it against whatever else has since raised its own activation in the meantime — which, in a phone-mediated life, is usually the next fifteen things the phone put in front of you.

This is not merely a theoretical claim. Altmann and Trafton tested the model against latency and error data from a task — the Tower of Hanoi — that depends heavily on suspending and resuming subgoals, and the interference and strengthening constraints they derived from the model fit that data.2 A decade of follow-up work then took the model out of the puzzle lab and into settings closer to ordinary life. Trafton, Altmann, Brock and Mintz showed that people who get a short warning before an interruption arrives spend that lag actively preparing — rehearsing where they were — and resume faster than people interrupted with no warning at all, exactly as a decay-and-rehearsal account predicts.3 Monk, Trafton and Boehm-Davis then varied how long an interruption lasted and how cognitively demanding it was, across three experiments on a hierarchical, interactive task, and found that both duration and demand independently lengthen the time it takes to resume the original goal — longer and harder detours cost more, in a way that tracks decay for short interruptions and tracks rehearsal opportunity as interruptions get longer.4

None of this is about phones specifically. It generalises to phones because a phone that makes you leave a conversation to open another app is manufacturing exactly the structure these experiments built on purpose: an interruption of measurable duration and demand, inserted between forming a goal and acting on it. Sophie Leroy's independent line of work gives this the name that has stuck outside the lab — attention residue — and shows the cost runs in the other direction too: people who switch to a second task with a first task left unfinished perform worse on the second task than people who either finished the first task or were forced to stop it under a hard deadline, because part of their attention stays locked on the interrupted goal rather than transferring cleanly.5 Gloria Mark's field studies, running independently of the Trafton and Leroy lines, measured the same territory from the outside: knowledge workers who get interrupted do eventually finish the interrupted task, and often just as accurately, but only by working measurably faster afterward and reporting more stress, frustration and time pressure in the process.6 The task gets done. It costs more to get it done. That is the same finding as the resumption-lag experiments, observed at the level of a whole workday instead of a single trial.

Put the three strands together and the definition sharpens past the everyday meaning of "switching cost." It is not merely "takes time to switch." It is: a suspended goal loses retrievable strength the longer and more demanding the detour away from it is, and resuming it exacts a further cost in speed or stress even when the goal is eventually retrieved. That is a mechanism with moving parts — decay, rehearsal, interference, residue — not a vibe. And it makes a specific, checkable prediction about interfaces: the cost of using a tool is not the time spent inside the tool. It is the size of the detour required to get to the tool in the first place.

4. Why "just open the app" is a bigger ask than it sounds

Apply §3's mechanism to the ordinary phrase "just open the app" and the phrase stops sounding trivial.

Opening a separate app for a self-directed task — log the meal, do the check-in, review the budget — is, in the vocabulary of §3, an interruption you impose on yourself: you suspend the goal that made you reach for the phone, search for an icon, wait for a load screen, re-orient to an unfamiliar layout, and only then re-encode the goal you started with. Every one of those steps is exactly the kind of demand Monk, Trafton and Boehm-Davis showed lengthens resumption time.4 None of it is long by the standard of a phone call or a meeting — the interruptions in the lab studies run from seconds to a couple of minutes — but the mechanism does not require a long interruption to bite. It requires only that the detour be nonzero and that something more urgent-looking than the original goal be available during it, which on a modern phone is guaranteed: the app store icon sits next to a badge count, the lock screen carries three unrelated notifications, and every one of them is competing for exactly the activation your half-formed goal needs to survive the trip.

This is the mechanism-level account of a pattern that is otherwise easy to describe and hard to explain: personal apps that require a deliberate, separate open have a well-documented tendency to be installed once and used briefly. Baumel, Muench, Edan and Kane analysed real-world Android usage panels across a large sample of mental-health apps — trackers, peer-support tools, meditation and breathing-exercise apps, the exact category my friend's four wellness apps belong to — and found a median daily open rate across the sample of about 4%, with wide variation by category and, on breathing-exercise apps specifically, a median 30-day retention of effectively zero.7 Industry-reported retention aggregates for consumer apps generally describe a similar shape at a coarser grain — a large majority of installs churning within roughly the first three months, with a meaningful share abandoned after a single use — figures that are secondhand and should be read as order-of-magnitude rather than as an audited rate, but that agree in direction with the panel-based finding.8

I want to be careful about what this does and does not show. It does not show that these apps failed because of the interruption mechanism specifically — poor onboarding, weak content, no reason to return, and simple lack of felt need are all live alternative explanations, and none of the studies cited isolates "requires opening a separate app" as a variable against an otherwise identical messaging-native version, because — as §1 says plainly — that study does not exist. What it shows is that the failure pattern §3 predicts (a goal that has to survive a detour tends not to survive it) is consistent with the failure pattern actually observed in the category of app this thesis is about. Consistency is not proof. It is the strongest thing available short of the experiment nobody has run, and it is the reason this paper's central claim is stated as a mechanism plus a pattern, not as a settled causal finding.

5. Messaging is not an interface choice, it is where the attention already is

The argument so far would already apply somewhere with a thriving app ecosystem and cheap mobile data. It gets sharper, not just more convenient, in the market this series writes for.

Start with what "opening an app" costs before any cognitive mechanism is even in play. Earlier papers in this series have made the case that African institutions are frequently building for a user whose device, data plan and attention are all more constrained than the ones assumed by software built elsewhere. GSMA's 2025 accounting of the region's mobile economy puts the "usage gap" — people covered by mobile broadband who are not yet using the mobile internet at all — at roughly 63% of the covered population.9 For someone in that gap, or just above it, a new app is not a free download; it is bytes against a data bundle and a slot on a phone that may already be short of storage. Messaging carries none of that marginal cost, because the messaging app is already open, already paid for, and already the thing the phone is mostly used for.

And in this market, one messaging app in particular is not a platform among several — it is closer to the substrate everything else runs on top of. WhatsApp is reported to reach roughly 95% of Kenya's internet users, with a 79.3% monthly-use share and around 20.85 million user accounts in the country;10 penetration figures reported for Nigeria and South Africa run in a similar 96–97% range of each country's internet users.11 These are platform-reported and press-aggregated figures rather than independently audited counts, and I use them the way this series uses this kind of number: as an order of magnitude, not a precise rate. But the order of magnitude is doing real argumentative work here. A near-universal messaging habit is not "an interface choice a designer made." It is closer to a fact about the environment a designer is building into, the way electricity supply or road quality is a fact about the environment — and Pielot's notification-behaviour study gives a reason beyond raw penetration to expect the habit to be strong: across the sample studied, messenger notifications were checked faster than any other category, out of a background of roughly 63.5 notifications received per day on average.12 People are not merely present in messaging. They are primed to attend to it faster than to anything else competing for the same attention.

Put §3 through §5 together and the shape of the argument is this. A tool that requires a separate open imposes, on every single use, the exact structure — a detour of nonzero duration and demand, with competing stimuli available during the gap — that the resumption-cost literature shows degrades a goal's chance of surviving to completion. A tool resident in the messaging thread does not impose that detour, in a market where the thread is already the place attention defaults to and is checked fastest. That is not a claim that messaging-native tools are more pleasant to use. It is a claim that they remove, structurally rather than through better design, the specific mechanism this paper has spent three sections establishing as the thing that kills personal-admin and self-improvement tools in practice.

6. The objection with teeth: conversational interfaces are bad at complex tasks

The most serious objection to everything above is not "but people like apps." It is that chat, as an interface, has a demonstrated failure mode with complex tasks, and no amount of removing switching cost fixes a mode of failure that switching cost was never the cause of.

Janssen, Grützner and Breitner's study of real-world chatbot deployments is the sharpest version of this evidence I found. Working from a sample of 103 real-world chatbots and twenty expert interviews, they derive twelve critical success factors and, separately, a set of concrete failure reasons — among them that chatbots are poorly suited to deeply complex processes, and that many deployments underestimate the technical, financial and organisational resources the ongoing maintenance of a conversational system actually requires.13 The paper reports having measured a discontinuation rate across its sample over a fifteen-month window; I was not able to obtain the specific figure from the sources available to me, and I record that as a gap rather than estimate it — see Appendix B. The qualitative failure reasons, which I can report with confidence, are enough to make the point: conversational interfaces fail in practice for reasons that have nothing to do with switching cost and everything to do with the shape of the task itself. A multi-step booking flow with branching options, a form with a dozen required fields, a comparison across many similar-looking products — these are tasks where a visual interface lets you see the whole state at once, and a chat interface makes you reconstruct that state one turn at a time, holding more of it in your own head than a screen would ask you to.

This objection has real force against a certain kind of Rafiki, and I do not think the honest response is to argue chat is secretly fine for everything. It is not. The response is that the objection is sharpest against exactly the category of task Rafiki's own product description does not claim to specialise in doing inside the chat window. "Rafiki... acts on your behalf in the real world — booking, ordering, reminding, following up"1 describes an assistant that delegates the complex multi-step interaction to whatever system actually handles it — a booking API, a merchant's checkout, a calendar — and keeps the user-facing surface to what messaging is actually good at: a short instruction in, a confirmation or a clarifying question back. The chat is the front door and the accountability layer, not the form. Where a task genuinely requires seeing many options at once — comparing five flights, reviewing a multi-line invoice — the honest design implication of this section is that the messaging surface should hand off to a link, a card, or a lightweight view rather than pretend the whole interaction belongs in a text bubble. That is a real design constraint this paper's own argument imposes on Rafiki, not a problem the argument gets to wave away.

7. The objection that should worry us most: this is also the most invasive surface available

Every property that makes messaging the right place to remove switching cost is also, without exception, a property that makes it the most dangerous place to put a persuasive AI.

The previous paper named this risk under the heading of honest risk and moved on. I want to sit in it longer, because it is not a separate concern sitting next to this paper's thesis — it follows from the same mechanism. §5 argued that residency in the thread works precisely because it removes friction: no detour, no separate open, no chance for the goal to decay before you act on it. Friction is exactly what stands between a persuasive system and unwanted influence. A tool a user has to deliberately open has, built into that requirement, a moment of reconsideration — Am I sure I want to do this right now? A tool that is simply present in the conversation, timed to arrive exactly when the user's guard is down and their goal is freshest, has removed that moment along with the switching cost, and there is no version of the mechanism in §3 that lets you keep one and discard the other. They are the same door.

Layer onto that the fact that "lives in your messages" means, literally, access to the most intimate record most people keep of themselves — who they are angry at, what they are ashamed of, what they promised someone and did not do. An assistant that can act on that record on the user's behalf is holding something closer to a diary with write access than a productivity tool. The commercial incentive of whoever operates it does not disappear because the product description says its "job isn't to obey" — a system built by a company still has the company's interests threaded through it somewhere, whether that shows up as which merchant a booking defaults to or which behaviour gets reinforced because it is measurable and monetisable rather than because it is what the user actually wanted. I do not have a mechanism-level fix for this the way §3 supplied one for the switching-cost problem, and I am not going to pretend I do. The honest position is that this paper's own argument for why messaging-native works is simultaneously the argument for why it should be trusted less by default than an app the user can simply delete, and any actual deployment has to be judged against that, not against how convenient it is.

8. Notification fatigue is real, and it does not save the objection — but it does not defeat this paper's claim either

The third serious objection is empirical rather than ethical: notification fatigue is documented and appears to be getting worse, not better, and a paper arguing that residing in someone's attention is good design should have to answer for adding to it.

Pielot's figure of roughly 63.5 notifications a day, mostly from messaging and email, checked within minutes regardless of whether the phone is silenced,12 already describes an attention economy under real pressure before Rafiki exists. Introduce a proactive assistant that checks in on the user unprompted — which is precisely what Rafiki's own description promises1 — and the naive worry is obvious: this is one more thing competing for a budget of attention that the cited research suggests is already strained, and "resident in your messages" could just as easily mean "one more source of the fatigue" as "the tool that finally sticks."

I take that worry seriously and do not think it is fully answerable in the abstract — it depends on implementation choices this paper has no data on, because there is no deployment (§1). What I can say is that it is a different claim from the one this paper is making, and worth separating cleanly. This paper's mechanism in §3 is not "more interruptions are good." It is "a detour between forming a goal and acting on it degrades the goal, so removing the detour removes a specific, measured cost." A badly built Rafiki that pings constantly and unpredictably reintroduces exactly the interruption-and-demand structure §3 describes, except now the interruption is coming from the tool that was supposed to remove it, and the mechanism this paper relies on would predict that version fails too, for the same reason the four wellness apps in the opening paragraph failed. Notification fatigue is not a counterexample to the thesis. It is what happens when a messaging-native tool is built as though residency alone were sufficient, without also respecting the same activation-and-attention budget the thesis is built on. That is a real design risk, and I register it as one of the conditions under which this argument would be shown wrong in practice (§10), not as an objection this paper gets to dismiss.

9. The objection I have to state plainly: Flux is building this

Rafiki is a Flux product, currently in research, and I am an author at Flux writing a paper that concludes the interface Flux has chosen is structurally superior to the alternative. That is a conflict of interest and I would rather name it in one sentence than let a reader find it themselves and wonder why I didn't.

The mitigation available to me is the same one this series has used before: does the paper's actual conclusion make Rafiki's job easier or harder? I think it makes it harder. §6 concludes that Rafiki's chat surface has to hand off multi-step, high-branching tasks to something other than the text window, which is a real engineering constraint, not a talking point. §7 concludes that the same mechanism making the product work is the mechanism making it the single most sensitive kind of product Flux could build, which argues for slower, more cautious deployment, not faster. §8 concludes that a naively built version of exactly this product would fail for a reason this paper itself supplies. None of that is the argument a paper writes to make a product sound easy to ship. If the honest incentive were to make the case as friction-free as possible, this is not the paper that would result.

10. What would falsify this

Stated as commitments, in decreasing order of how much damage each would do to the claim.

A controlled comparison — the study §1 says does not yet exist — finds that a messaging-native version of a well-specified personal task performs no better, or worse, than an otherwise identical standalone-app version on completion or long-run retention. That would show the switching-cost mechanism is not, in fact, the binding constraint this paper claims it is, and I would withdraw the central claim rather than the details around it.

Standalone personal-admin or self-improvement apps that use strong scheduled notifications and low-friction reopening (a single tap from the lock screen, no navigation) achieve retention comparable to what this paper predicts for a messaging-native tool. If the detour can be engineered away inside an app just as effectively as by moving the tool into a thread, then residency in messaging is not doing unique work — good notification design is doing the work, and messaging is just one way to get it.

Rafiki ships and shows the same abandonment curve as a typical standalone app despite being resident in the thread. This would not falsify the cognitive mechanism in §3, which is well established independently of this product, but it would falsify this paper's application of that mechanism to Rafiki specifically, and would mean some other, unaccounted-for factor is dominating the outcome.

Users in the markets this paper describes report that the messaging surface itself feels like more effort to use for this purpose than a dedicated app would — that the thread is experienced as cluttered or socially loaded in a way a private app is not. That would invert a load-bearing assumption of §5, that residency reduces rather than adds to the user's felt cost of engaging the tool, and would need to be taken as seriously as the mechanism it undermines.

11. Limitations, and what this paper does not claim

It does not report results from Rafiki. There is no deployment, and every sentence in this paper that mentions Rafiki by name is a prediction from the argument, not an observation of the product.

It does not claim to have found or run the controlled experiment that would isolate interface choice as a causal variable, holding the underlying task and content constant. That experiment does not appear to exist in the published literature, and this paper's contribution is connecting a well-established mechanism (§3) to an observed pattern (§4) — an inference, clearly weaker than a direct test, and Appendix B carries the gap forward explicitly rather than letting it disappear into the prose.

It does not claim the interruption-and-resumption literature was built with phones or messaging apps in mind. Altmann and Trafton's model was tested against the Tower of Hanoi; Monk, Trafton and Boehm-Davis's experiments used a synthetic interactive task; Mark's field studies were conducted on office knowledge workers switching between windows on a desktop screen, not on phone-app switching in a mobile-first African market specifically. Applying the mechanism to app-switching on a phone is a generalisation I believe is reasonable — the underlying claim is about goal activation and decay, not about any particular device — but it is a generalisation, and a reader should treat it as one rather than as a direct finding.

It does not claim messaging-native design is sufficient for a personal AI to be good, safe, or worth trusting. §7 and §8 are explicit that the same mechanism which makes the interface effective also makes it more dangerous and, badly implemented, no better than the alternative it is meant to replace. Removing the switching cost is a necessary-looking condition for this category of tool to get used at all. It is nowhere near a sufficient one for it to deserve being used.

And it does not resolve the privacy and power questions raised in §7. Naming a risk is not mitigating it, and a future paper — ideally one written before rather than after Rafiki ships — owes the series an actual account of what limits on data retention, persuasion design and default-on behaviour would make the risk in §7 tolerable rather than merely disclosed.


Appendix A — Evidence table

S peer-reviewed source, P press-reported or industry aggregate, I internal Flux product record, not sourced.

# Claim used in this paper Value Source
1 Memory-for-goals model; goal activation decays after suspension, is restored by rehearsal and cues qualitative model, tested against Tower of Hanoi latency/error data S2
2 Warned interruptions produce more preparatory rehearsal and faster resumption than unwarned ones directional finding, controlled experiment S3
3 Interruption duration and demand each independently lengthen resumption time directional finding across three experiments S4
4 Attention residue: unfinished-task switching degrades performance on the subsequent task directional finding, controlled experiments S5
5 Interrupted tasks are completed at similar accuracy but with more speed, stress and frustration field study, knowledge workers S6
6 Median daily open rate across sampled mental-health apps ~4% (IQR 4.7%), varies sharply by category S7
7 30-day retention, breathing-exercise apps in the same sample ~0% median S7
8 Consumer app abandonment within ~90 days roughly 70–80% of installs (industry aggregate) P8
9 Mobile broadband "usage gap," Sub-Saharan Africa ~63% of covered population not yet using mobile internet P9
10 WhatsApp share of Kenyan internet users ~95%, 79.3% monthly-use share, ~20.85M accounts P10
11 WhatsApp share of Nigerian / South African internet users ~96–97% each (press-aggregated) P11
12 Average daily mobile notifications; messenger notifications checked fastest ~63.5/day average S12
13 Chatbot failure reasons include unsuitability for deeply complex processes and underestimated maintenance burden qualitative findings, 103 real-world chatbots + 20 expert interviews S13
14 Chatbot discontinuation rate over the 15-month study window
15 Head-to-head completion/retention comparison, messaging-native vs. standalone-app delivery of an identical personal task
16 Rafiki deployment retention or completion data — (product in_research)
17 Felt effort of engaging a personal-AI task inside an existing messaging thread vs. inside a dedicated app, from users in an African market

Rows 14–17 are the measurements that would move this paper from an argument to a finding. Row 15 is the one I would fund first: it is the direct test of the paper's central inference, and nothing else in Appendix A substitutes for it.

Appendix B — Open questions and citation register

CITATION-NEEDED — the direct comparison. No controlled study was found comparing completion or retention for an identical personal task delivered messaging-natively versus as a standalone app. This is the single largest gap in the paper's evidentiary basis, named plainly in §1 and again in §11, and it is the reason the paper's claim is stated as a mechanism-plus-pattern inference rather than a causal finding.

CITATION-NEEDED — the exact chatbot discontinuation rate. Janssen, Grützner and Breitner report analysing 103 real-world chatbots to examine a discontinuation rate over 15 months.13 The specific percentage was not obtainable from the sources available during this paper's research and is not stated anywhere in this paper as a number. Appendix A row 14.

CITATION-NEEDED — primary telecom or regulator data for WhatsApp penetration. §5's figures for Kenya, Nigeria and South Africa are drawn from DataReportal/Kepios's platform-reported aggregates rather than from a national regulator or from Meta's own disclosed per-country numbers. They are directionally consistent across three independent country reports and with this series' general finding that WhatsApp dominates African messaging, but should be checked against primary telecom-regulator data before being quoted as anything more precise than an order of magnitude.

CITATION-NEEDED — generalising the resumption-cost literature to mobile app-switching. The foundational experiments (§3) were run on puzzle tasks and desktop window-switching, not on phone-app switching. I have not found a study that runs the Altmann-and-Trafton or Monk-Trafton-Boehm-Davis paradigm on smartphone app-switching specifically. If one exists, it would be the strongest available direct test of this paper's core generalisation and I have not located it.

CITATION-NEEDED — Rafiki-specific evidence. Everything in this paper about Rafiki is inference from the argument, not observation of the product, because the product has not shipped. Any future revision of this paper should replace every predictive sentence about Rafiki with an observed one, as soon as there is something to observe.

  1. Flux internal product record (Product model, slug rafiki, status in_research). Quoted once, in full, in §1 and referenced in §6 and §8: "Rafiki is the personal operating layer: an AI that lives in your messages, not in an app you have to remember to open. It talks the way you talk to a close friend, acts on your behalf in the real world — booking, ordering, reminding, following up — and checks in on you proactively, the way someone genuinely in your corner does. Its job isn't to obey; it's to hold you to the person you said you wanted to be." Not a published source; treated throughout as a statement of design intent, not as evidence of outcomes.

  2. Altmann, E. M., & Trafton, J. G. (2002). Memory for goals: An activation-based model. Cognitive Science, 26(1), 39–83. https://doi.org/10.1207/s15516709cog2601_2

  3. Trafton, J. G., Altmann, E. M., Brock, D. P., & Mintz, F. E. (2003). Preparing to resume an interrupted task: Effects of prospective goal encoding and retrospective rehearsal. International Journal of Human-Computer Studies, 58(5), 583–603. https://doi.org/10.1016/S1071-5819(03)00023-5

  4. Monk, C. A., Trafton, J. G., & Boehm-Davis, D. A. (2008). The effect of interruption duration and demand on resuming suspended goals. Journal of Experimental Psychology: Applied, 14(4), 299–313. https://doi.org/10.1037/a0014402

  5. Leroy, S. (2009). Why is it so hard to do my work? The challenge of attention residue when switching between work tasks. Organizational Behavior and Human Decision Processes, 109(2), 168–181. https://doi.org/10.1016/j.obhdp.2009.04.002

  6. Mark, G., Gudith, D., & Klocke, U. (2008). The cost of interrupted work: More speed and stress. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI '08), 107–110. ACM. https://doi.org/10.1145/1357054.1357072

  7. Baumel, A., Muench, F., Edan, S., & Kane, J. M. (2019). Objective user engagement with mental health apps: Systematic search and panel-based usage analysis. Journal of Medical Internet Research, 21(9), e14567. https://doi.org/10.2196/14567

  8. Industry retention benchmarks aggregating app-analytics data (e.g., Business of Apps and Localytics-derived summaries, various report years), reporting roughly 70–80% of installed consumer apps abandoned within approximately 90 days, and a meaningful minority abandoned after a single use. Secondhand industry aggregation rather than a peer-reviewed measurement; used here strictly as order-of-magnitude context, not as a precise or audited rate.

  9. GSMA Intelligence, The Mobile Economy Africa 2025 (GSMA, 2025). Reports the mobile-broadband "usage gap" — population within network coverage but not using the mobile internet — at approximately 63% across the region.

  10. DataReportal / Kepios, Digital 2026: Kenya (2026). WhatsApp reported at approximately 95% of Kenyan internet users, a 79.3% monthly-use share, and approximately 20.85 million user accounts. Platform-reported and press-aggregated; order-of-magnitude, not audited.

  11. DataReportal / Kepios, Digital 2026: Nigeria (2026) and Digital 2026: South Africa (2026). WhatsApp reported at approximately 96–97% of each country's internet users. Platform-reported and press-aggregated; order-of-magnitude, not audited.

  12. Pielot, M., Church, K., & de Oliveira, R. (2014). An in-situ study of mobile phone notifications. In Proceedings of the 16th International Conference on Human-Computer Interaction with Mobile Devices and Services (MobileHCI '14), 233–242. ACM. https://doi.org/10.1145/2628363.2628364. Reports approximately 63.5 notifications received per day on average, predominantly from messaging and email, with messenger and social-network notifications checked fastest among all categories.

  13. Janssen, A., Grützner, L., & Breitner, M. H. (2021). Why do chatbots fail? A critical success factors analysis. In Proceedings of the 42nd International Conference on Information Systems (ICIS 2021), Austin. AIS Electronic Library. https://aisel.aisnet.org/icis2021/hci_robot/hci_robot/6. Analyses 103 real-world chatbots and 20 expert interviews to derive 12 critical success factors and a set of failure reasons, including unsuitability for deeply complex processes and underestimated ongoing maintenance burden. The paper's own discontinuation-rate figure over its 15-month study window was not obtainable from the sources available and is not quoted here — see Appendix B.

Ken Ruto
About the author
Ken Ruto

Founder of Flux. Building vertical AI-powered SaaS for Africa's institutions — and writing the thesis behind every bet. kenruto.fluximpact.org →

Share X LinkedIn WhatsApp
Did this land?
Was it useful?

Comments

No comments yet — be the first.

Replying to · cancel
Get new essays

No spam — just the next piece when it's out.

Think I got something wrong? Highlight any sentence to push back on it — or It comes straight to me, never shown publicly.

Push back
Related writing
31 min
The Word "Personalized" Is Doing Two Different Jobs
The skincare industry's "personalization" is mostly segmentation by skin type and budget, not the evidenced kind melanin-rich skin needs — personalization against post-inflammatory hyperpigmentation, a risk the mainstream clinical and testing base was never built to price in.
9 min
The Second Question: What Changes When an Explainer Can Be Asked Again
An inline explainer that answers once is a dictionary. The follow-up carries the value — and it is why the thing has to go standalone.
9 min
The Panel Is Also a Tab
Every AI answer surface solves the new-tab reflex by moving you somewhere smaller. The cost was never the tab — it was getting back.