Back to Work
Case Study 06 · AI-Native

Yulu — an AI-native micro-mobility product.

The best AI in this product is the AI you never see. Designed from a real research base. Four narrow agents that prevent failures instead of apologising for them.

ContextDelhi Micro-Mobility
RoleAI Design Engineer
ScopeEvidence → Prompts → Prototype
Year2026
9:41 To Connaught Place · 6.2 km YULU ZONE YULU ZONE 24 km 19 km 11 km 3 bikes can make this trip Needs 9.3 km of range · 2 bikes hidden Yulu 4471 120 m · 4 min walk 24.0 km range Yulu 6135 340 m · 9 min walk 19.5 km range
23 rider interviews in the evidence base this was designed from — real research, conducted in Delhi in 2026.
4 narrow agents in the architecture. Only one of them is a language model — and that restraint is the design decision.
3 system-prompt versions before the billing agent behaved. V1 invented refunds that didn't exist.

Most products add AI at the end. This one couldn't have been.

The Problem

Riders weren't leaving over price or roads. They were leaving because the system surprised them — locks that lie, batteries that die mid-route, charges with no evidence.

My Approach

Killed the obvious chatbot. Designed four narrow agents that resolve failures before the rider meets them, and stay silent otherwise.

A chatbot is a well-designed apology for a failure that should have been prevented.

What is real here, and what is a design exploration

Real: the 2026 research base — system study, field observation, 23 rider interviews, personas, journey maps, SWOT and gap analysis — by a team of three (Aalvee Damle, Bhanuja Dass, Vrinda Singh) for a service design course.

A design exploration: everything after it is mine and was never shipped to a live fleet. Evaluation results come from scripted scenarios — directional, not empirical.

I started with evidence, not with a model.

Riders had made peace with price and roads. They hadn't made peace with a system that behaved unpredictably.

23personal interviews, mostly first-time riders — the brief was deliberately "staying a beginner".
2Yulu zones observed in the field: INA and Mandi House metro stations, including conversations with ground staff.
4personas and 2 full customer journey maps, spanning past experience through post-ride.
26SWOT entries and 6 named gaps — the raw material the synthesis had to work across.
THE 2026 EVIDENCE BASE · SIX METHOD STREAMS System study Input · throughput output · feedback environment Surfaced: locking, abandonment, battery Field observation INA + Mandi House Yulu zones Ground-staff talks Surfaced: cash at the bike, staff dependence 23 interviews 15-question guide + open conversation Beginner's lens Surfaced: app friction, battery anxiety, cost Persona mapping 4 personas, 21–25 Students + MNC South / North Delhi Surfaced: occasional vs daily usage split Journey maps 8 stages, past experience → after usage Surfaced: satisfaction splits after the ride SWOT + gaps 11 S · 7 W · 5 O · 3 T 6 named gaps Surfaced: weak tracking, theft, infrastructure One corpus · ~14,000 words of research artefacts Slides, transcript notes, maps and matrices — normalised into plain text before anything else happened The research team's recommendation Eliminate the ₹250 security deposit. Let one-time riders pay cash to ground staff before the ride, settle the balance by time afterwards, and reward members with points. Correct — but already table stakes in 2026, which is where this project starts.
FIG 01The research dataset. Six method streams, normalised into one plain-text corpus so a language model could read all of it at once. The research team's own recommendation is included deliberately — the synthesis had to argue past it.

Complaints split into two kinds

The journey map's after-ride complaints look like six equal problems. Three — roads, infrastructure, storage — are outside the product's control. The other three are reliability failures the product causes.

Price had already been accepted

The personas describe riders as "price-conscious but not a constraint". Unpredictability was the one thing nobody had accepted.

I used AI to read the research against me, not for me.

I gave the model the whole corpus at once, and briefed it to make agreement expensive.

ProductDesignAIEngineering

Kept — the reframe

"App could be smooth", "weak tracking", "poor locking" and "payment doesn't go through" are one failure in four costumes: the rider can't tell what state the system is in.

Rewritten — the money cluster

The model caught that the team's own fix adds a cash step the research calls friction. But it blamed ground staff, who actually rescue failed rides. I reframed it as charges the rider can't evidence.

Rejected — "safety on the road"

A plausible fifth cluster built from one journey-map cell, with no interview behind it. Proof that model output is a proposal, not a finding.

Kept — the thinness warning

All 23 participants were beginners, so the data can't support retention claims. Every retention claim below is framed as a hypothesis with a test.

OPPORTUNITY MAP · FREQUENCY × TRUST DAMAGE Mentioned once or twice Mentioned across most of the corpus FREQUENCY OF MENTION → TRUST DAMAGE → Stops using Complains WHERE THE PRODUCT GOES Lock state can't be trusted System study · 4 SWOT weaknesses Battery dies mid-route Every journey-map stage from comparison onward Charges you can't evidence Deposit · cash at the bike · disputed post-ride time QR / GPS won't connect Brainstorm current-problems column · interviews Road quality & infrastructure High frequency · outside product control Price "Price-conscious but not a constraint" — already accepted Storage space on the bike Gap list only · hardware decision Parking / zone density Ops footprint, not interface "Safety on the road" — rejected Model-generated from one journey-map cell. No source behind it.
FIG 02The opportunity map. Plotting by trust damage separates the top-right cluster from what riders had already accepted about Delhi. The dashed red point is the rejected finding, kept on purpose to show where judgement was applied.
The problem was never supply or price. It was invisible failures eroding trust — and the rider carrying the cost of the system's uncertainty.

The obvious AI product here was a chatbot. I killed it.

Four directions, scored against criteria written before I saw them. The chatbot had the highest demo value and the lowest product value.

ProductDesignAIEngineering
Product hypotheses
Four directions on the table
3 mine · 1 model-generated
1 chosen
H1 · MINE

A ride-hailing-grade trust layer

Show real bike state before she commits: verified lock, live battery, confirmed availability.

Necessary, not sufficient
H2 · MINE

Deposit-free, wallet-based billing

No ₹250 deposit, no cash, auto-settle from a wallet. The research team's recommendation, fully digital.

Table stakes in 2026
H3 · THE EXPECTED ONE

A rider-facing AI concierge

A conversational assistant for unlock failures, disputes and route questions.

Rejected
H4 · CHOSEN

AI as invisible connective tissue — narrow agents that prevent failures rather than discuss them

No chat surface and no score shown to the rider. Single-purpose agents decide what appears on the map, whether this bike can make this trip, and how a dispute resolves — surfacing only when a human should make the call.

Chosen — designed and prototyped below
Decision matrix
Scored 1–5 against criteria written before the options
Weighted toward evidence and failure cost,
not novelty
Direction Evidence
support
Needs AI
to work
Cost of
being wrong
User
control
Buildable
as a proto
Result
H1 Trust layer51445Folded into H4
H2 Wallet billing51345Folded into H4
H3 AI concierge25234Rejected
H5 Crowd-sourced health
Model-generated
32235Rejected
H4 Invisible agents54543Chosen

Why the chatbot died

  • Ananya doesn't want a better way to report a broken bike. She wants to not get one.
  • Every chat turn exists because the system already failed.

The model's idea I rejected

  • Crowd-sourced bike health: riders rate each bike after the ride.
  • Survivor-biased — the rider with a jammed lock is the least likely to rate.

The column that decided it

  • "Needs AI to work": remove the intelligence — does the product still keep its promise?
  • Only H4 needs per-trip judgement over live signals at fleet scale.

Design for the rider who is two bad rides from leaving.

I designed for the daily rider — and named the rider that choice leaves behind.

Ananya Verma, 29
Composite of the Office Goer and Daily Life Traveller archetypes · South Delhi · rides 3–4 times a day
Trip
~4 km from the metro to her office, plus stops on the way home. Six months in.
What burned her
A bike that died 2 km from the office, and a lock that read locked but wasn't — followed by a billing dispute nobody could evidence.
What she needs
The app to stop surprising her.
How she leaves
Quietly. An auto instead, three days running, and then it's a habit.
The counter-persona I'm not designing for
Named so the trade-off is examined, not convenient
Who
The 21–24 occasional, price-sensitive student from the research personas.
How she's still served
No deposit, payment at the bike, and a failed unlock never charges her.
Open
Whether an empty map reads as diligence or as a broken app to her. I'd test it first.

AI generated. I decided.

AI did the volume, pattern-finding and first drafts. Every decision about what the product is stayed with me.

Me · the designer

  • Defined the problem — six complaints reframed as one trust failure.
  • Set the constraints — no visible AI surface, no score shown, no failure resolved in the company's favour.
  • Designed the interaction — absence-based map, destination-first flow, payment at the bike.
  • Specified AI behaviour — asymmetry, escalation, message style, and what each agent isn't told.
  • Judged quality — wrote the rubric and reviewed every piece of generated code.

AI · in my design and engineering environment

  • Read the whole corpus in one pass and clustered it by cause.
  • Found the contradiction in the research team's own fix.
  • Generated 31 edge cases — I kept 9; one changed the architecture.
  • Wrote the prototype — state machine, screens, fault injection.
  • Flagged thin evidence — which limited what I'm allowed to claim.
HUMAN / AI RESPONSIBILITY MAP ME AI Frames the problem Rejects the weak cluster Clusters the corpus Flags thin evidence RESEARCH Picks the direction Writes the criteria first Argues the opposite Generates alternatives DIRECTION Designs the interaction Decides when AI is visible Drafts microcopy 3 variants per state DESIGN Writes the behaviour spec Sets the asymmetry rule Drafts prompt language Fails 12/12 on v1 AI BEHAVIOUR Reviews + fixes code Catches the render bug Writes the prototype State machine, screens BUILD Scores it Generates edge cases EVAL
FIG 03The responsibility map. Model output enters as a proposal and only leaves as a decision after passing through the top row — nothing the model produced shipped unreviewed.

Four narrow agents. Only one of them is a language model.

Most of this system is thresholds and a decay curve — because most of it should be.

ProductDesignAIEngineering

Bike Readiness

Deterministic · rules
Job
Decide whether a bike appears at all — for this trip — and record why.
Why not a model
Hard signals with clear thresholds. A model would add latency and unpredictability to a question with a correct answer.

Battery Decay

Predictive · regression
Job
Predict whether this bike dies mid-route on this specific trip.
Why not a model
Battery decay is physics with a fitted curve. A language model doing arithmetic is the wrong tool in a fashionable hat.

Billing Resolution

LLM · reasons over evidence
Job
Resolve lock and battery disputes without a human — adjust, take no action, or escalate — and explain it in one sentence.
Why a model
The one true judgement: incomplete, sometimes contradictory evidence, open-ended failure modes, and an answer written for a human.

Rebalancing / Dispatch

Ranking + LLM rationale
Job
Predict zone supply 1–3 hours out and give ground staff a ranked task list, each with a plain-language reason.
Human in the loop
Staff can override, and overrides feed back into readiness confidence.
AI PRODUCT ARCHITECTURE SIGNALS Lock-state ping GPS trace + drift Charge + cycle count Service history Destination + distance Fare ledger Zone demand history AGENTS Bike Readiness required = max(trip × 1.5, 2 km) + lock · GPS · service thresholds RULES Battery Decay Fitted decay curve per battery age → predicted range for this trip REGRESSION Billing Resolution Reasons over ride telemetry adjust · no action · escalate LLM Rebalancing / Dispatch Ranked task list · LLM writes the reason each task exists lock confidence lowers range trust range flags pre-attach to the ride readiness flags set pickup priority SURFACES Rider app Map shows only cleared bikes — no score, no badge, no greyed-out pin. Agents surface only when confidence is low enough that a human should decide. Silent — the default state Surfaced — 3 of 17 screens Marginal-trip notice · empty-map explanation · billing adjustment. Everything else is the agents acting unseen. Ops dashboard Opens to a ranked task list, not a raw map — "what to do next", not "everything happening". Every task states why it was generated. Staff can override, and the override feeds back into Readiness confidence. The feedback loop that keeps this honest Staff overrides are the only measurement of how often the system hides a bike that was actually fine. It is designed in, not bolted on.
FIG 04The architecture. The lime arrows make this orchestration rather than four features: agents pass confidence and flags to one another, so billing never re-investigates a dispute from scratch. Remove one agent and the other three degrade.
The other half of the architecture
What is deliberately not AI
A design decision,
not an omission
CapabilityImplementationWhy it stays dumb
PricingFixed rate tableA rider must be able to predict her fare before she commits.
Zone geofencingPolygon containmentA bike is in a zone or it isn't. A probabilistic answer to "can I end my ride" would be cruel.
Payment stateExplicit state machineMoney movement should be auditable step by step, not interrogated.
Fare stop on failureEvent-triggered ruleBattery death or an unconfirmed lock stops the meter instantly. Refunds are judgement; stopping the charge isn't.
The rule I applied

A capability earns a model when its input is incomplete, its output open-ended, and a wrong answer recoverable. Billing meets all three. Pricing, geofencing and payments meet none.

The central question wasn't what the AI knows. It was when the AI should speak.

Silence is the default. The AI speaks only when a human needs to make the call.

ProductDesignAIEngineering
Interaction treatments
How should bike readiness appear on the map?
Same agent output.
Three ways of handing it over.
TREATMENT A
94% 71% 38%

Numeric confidence

Precise, but it hands the decision back: she has to turn "71%" into "will this get me to work" in traffic.

Rejected — offloads the judgement
TREATMENT B

Traffic-light badge

Fast to read, but amber is an invitation. The late rider takes it and walks toward a predicted failure.

Rejected — invites the gamble
TREATMENT C · CHOSEN
24 km 19 km — nothing here —

Absence

A bike that can't make this trip isn't on the map. The mental model is one sentence: if it's here, it can get me there.

Chosen

Readiness depends on the trip

A bike only shows if its predicted range covers 1.5× the trip (never under 2 km) and its lock, GPS and service checks pass.

So the destination comes first

Readiness recomputes whenever the destination changes, so search became the entry point of the whole ride.

Absence has to explain itself

An empty map names the check — "none have the range for 6.2 km" — and the sheet narrates removals, so a vanishing pin reads as a decision, not a glitch.

The badge wasn't discarded, it moved

It reappears only in the marginal-trip notice. Three of seventeen screens show an agent speaking; the other fourteen are agents working.

USER FLOW — HAPPY PATH AND ITS FIVE FAILURE BRANCHES Destination first, not optional Map cleared bikes only Preview range vs. requirement Pay at the bike hold, not capture Unlock timer waits for the lock Ride fare always visible Park + lock system confirms, not you Receipt adjusted before you ask FAILURE BRANCHES — DESIGNED AT THE SAME FIDELITY AS THE HAPPY PATH F1 · NO BIKE QUALIFIES Empty map names the check out loud. Offers a shorter trip. Readiness · surfaced by absence F2 · UNLOCK FAILS Hold released, stated explicitly. Bike drops off the map for all. Retry capped at two F3 · BATTERY DIES The prediction was wrong. Fare stopped at death, not at end. Adjustment pre-applied F4 · EVIDENCE AMBIGUOUS GPS gap across the window. Not charged while it's reviewed. Billing · escalated to a human F5 · LOCK WON'T CLOSE "End ride anyway" stops the fare immediately. Bike flagged. Ananya's original bad ride, fixed Every branch resolves in the rider's favour by default. Not generosity — it is the behavioural rule from the specification expressed as interface: prefer a false negative over a false positive, and never resolve ambiguity in the company's favour. Hiding a working bike costs a rider forty seconds. Showing a broken one costs the relationship.
FIG 05The user flow. Seventeen screens, five of them failure states drawn at full fidelity — because trust is built or destroyed in exactly those five.

A behaviour spec first. Then the prompt that enforces it.

The spec is the design document. The prompt is its implementation — and it took three versions to match.

ProductDesignAIEngineering
AI behaviour specification
What each agent may do, and may never do
The contract the prompts implement
and the rubric scores against
AgentMay doMay never do
Bike ReadinessHide a bike. Flag it for pickup. Block ride start until the lock opens.Show a bike it isn't confident about. Hide one without a reason.
Battery DecayHide bikes below requirement. Raise a marginal-trip notice.Block the ride — "Continue" is always available.
Billing Resolution
the LLM
Lower a fare. Escalate to a human. Write one sentence to the rider.Raise a fare. Cite anything not in the ride data. Offer anything not on its action list.
DispatchRank staff tasks, each with a reason.Act without a human, or ignore a staff override.
Rule 1
Asymmetry

Hiding a working bike costs 40 seconds. Showing a broken one costs the relationship.

Rule 2
Uncertainty favours the rider

If the system can't prove the rider owes it, she doesn't.

Rule 3
Reasons in plain language

Every automated decision explains itself the moment it becomes visible.

Prompt iterations
Billing Resolution Agent · V1 → V3
Scored against 12 scripted scenarios
V1 · 0 / 12
"Resolve the dispute fairly."

Fluent and unusable. Invented discounts and callbacks the product doesn't have, asserted faults the data didn't show, and often decided nothing.

V2 · 8 / 12
Fixed actions + cite the evidence

Invented remedies disappeared. But under ambiguity it quietly sided with the company, and wrote timestamps at riders.

V3 · 11 / 12
Asymmetry, escalation, style rules

Ambiguity escalates, fares only go down, and messages are two plain sentences in clock time — no apology.

The one V3 still fails — and why I kept it

On one clean ride, V3 escalated instead of closing the case. I accepted it: over-escalating costs a reviewer a few minutes; under-escalating charges riders for failures nobody could prove.

Context engineering
What the billing agent is told — and what it is deliberately denied
The denials are the design

In context

  • The ride's event log — GPS, lock, battery and fare events, each citable.
  • The fare ledger for this ride only.
  • Flags from the other agents, so no dispute starts from scratch.

Withheld on purpose

  • Ride history and lifetime value — that's fairness priced by loyalty.
  • Prior dispute count — it punishes people for past bad luck.
  • Wallet balance and fleet margins — no judgement should depend on either.
The section I'd defend first in a review

An agent that cannot see who you are cannot treat you differently for being you. Enforced at the context layer, where a later prompt tweak can't quietly undo it.

Seventeen screens, five of them failures, all drawn at the same fidelity.

Grayscale first, so nothing was decided by colour — and failure states designed alongside the happy path, not after it.

ProductDesignAIEngineering
9:41yulu Where are you going? needs 9.3 needs 4.7 needs 14.1
A5 · Destination first
Required range shown before she commits
9:41yulu Change 24 km 19 km 11 km 24.0 km 19.5 km
A6 · Map
Failing bikes don't render at all
9:41yulu Change Notify me when one's ready Try a shorter trip Search another area
F1 · Empty map
Names the check so absence reads as care
yulu-wireframes.html
Open the annotated wireframes →
17 frames, 23 annotations. Every frame is tagged silent, surfaced or absent, so the invisible AI layer is legible at wireframe stage.
Design system
The components, and the rules they encode
Small on purpose — a product
this quiet doesn't need many parts
ComponentSpecThe decision inside it
Notice, four variantsWarning · error · success · infoColour always pairs with an icon and a bold heading. Meaning never rides on colour alone.
The list is not secondaryBottom sheet mirrors every map pinScreen-reader users get the full core task, not a reduced version of it.
Manual bike-ID entryFirst-class, not a fallbackScanning a QR code needs visual targeting and steady motor control that not every rider has.
Throttled live updatesAnnounced once a minuteFare and range change constantly; announcing every tick floods a screen reader.

A prototype you can break on purpose.

Working code, not a clickable flat — toggle a fault mid-ride and the failure states fire for real.

ProductDesignAIEngineering
yulu-prototype.html · live
Open full screen →
Try this: pick Nehru Place (9.4 km) and watch bikes disappear from the map that were there for a 3 km trip — the readiness rule recomputing against the new requirement. Then toggle "Lock won't confirm at end" and finish a ride: that is Ananya's original bad experience, redesigned so the cost lands on the system instead of on her.

Correction · readiness returns reasons, not yes/no

The inspector, ops dashboard and empty-map screen all have to explain why a bike is hidden. A model can't know a screen three steps away dictates what a function returns.

Correction · faults flip mid-ride

The model set failures before the ride started. Live toggles let a reviewer kill the battery during a ride and watch the fare stop — the demo of the failure became the argument for it.

You can't ship judgement you haven't scored.

The rubric came before the prompts, so "it reads well" could never be the standard. Scripted scenarios against my own prototype — not user or fleet data.

Evaluation rubric
Billing Resolution Agent · V3 across 12 scenarios
Any dimension below 3
blocks the version
Evidence groundingEvery claim cites an event present in the telemetry. No inferred mechanical causes.
12 / 12
Asymmetry complianceAmbiguity resolves to escalate, never to "no action". Uncertainty never favours the company.
12 / 12
Bounded remediesNo credits, discounts, callbacks or anything the product doesn't have.
12 / 12
Message qualityTwo sentences, no apology, clock time not timestamps, no component ever named.
11 / 12
Escalation correctnessEscalates when it should — and doesn't when it shouldn't.
11 / 12
DeterminismSame telemetry, same action across three runs. Wording may vary; the decision may not.
12 / 12
Edge-case matrix
31 generated · 9 kept · 1 changed the architecture
Generated by the model,
triaged by me
Edge caseSourceSystem responseRider costVerdict
Rider's phone dies mid-ride Model The bike's lock ends the ride; the fare stops at last confirmed movement. None Kept — changed the architecture. The bike, not the app, is the source of truth for ride end.
Two riders scan the same bike Model First confirmed unlock wins; the second sees the unlock-failure screen, hold released. Seconds Kept. Reuses an existing state.
Ride ends outside every zone Mine End blocked, nearest zone shown, any fee stated before she commits. A walk Kept. A fee is never discovered after the fact.
Rider spoofs GPS Model Rejected. A different project, and it would pull the product toward suspicion of the user.
Fleet-wide outage Model Rejected. Infrastructure, not product behaviour.
Rider is under age Model Rejected here. Real compliance work, unrelated to the trust thesis.
What 22 rejections say about working this way

The value isn't in the generation or the filtering — it's in having a written scope sharp enough to filter against.

How the hypothesis would be tested

The interesting failures weren't in one layer. They were between them.

Three failures stayed in one layer. The fourth lived in the gaps between design, spec and code.

01
Broke in: design · Fixed in: design

Destination-first stranded anyone without a destination

A rider who just wanted to browse had nowhere to go. Fixed with a browsing mode that applies only the 2 km minimum — and says so, never passing a basic check off as a trip check.

02
Broke in: the prompt · Fixed in: the prompt

V2 resolved ambiguity in the company's favour

The spec forbade it; the prompt never said so. Fixed by making ambiguity escalate explicitly. A value that isn't in the prompt isn't in the product.

03
Broke in: code · Fixed in: code

The map contradicted its own promise for a frame

Generated code drew every bike, then removed the failing ones — a one-frame flash of bikes that weren't real. Fixed by filtering before drawing.

04
Broke in: all three at once · Fixed in: design + spec + code

A designed screen that no rider could ever reach

The marginal-trip notice — the one place an agent speaks up — never fired. The spec gave no number, the code picked 1.2×, and readiness already required 1.5×, so no bike could ever qualify.

The fix touched all three: 1.6× in the code, the number written into the spec, and the calculation shown in the inspector.

THE WORKFLOW — WHERE THE LOOP ACTUALLY RUNS Evidence 2026 corpus Synthesis AI reads · I reject Direction Matrix · I decide THE LOOP — NOTHING HERE HAPPENS ONCE Design Flow · screens · states ME Behaviour + prompt Spec mine · draft AI's ME · AI DRAFTS Code AI writes · I correct AI · I REVIEW Evaluate Rubric mine · scenarios run a failed eval sends you back and often to more than one of them The thing conventional design process gets wrong about this work: these three are not sequential phases. A failing evaluation can send you to the interface, the prompt, or the code — and failure 04 above sent me to all three at once, because the bug lived in the gaps between them.
FIG 06The workflow. The top row happened once; everything inside the dashed loop ran repeatedly — and where a failed evaluation points back to is the diagnosis.

Three moments where the whole argument lands.

One screen where the AI is invisible, one where it speaks, and one where it admits it was wrong.

SILENT
To Connaught Place · 6.2 km 24 km 19 km 11 km 3 bikes can make this trip Needs 9.3 km of range · 2 bikes hidden Yulu 4471 · 120 m 24.0 km

The map, saying nothing

Three agents produced this screen; none appear on it. Two hidden bikes are noted in one line, never as a score.

SURFACED
← Back to map Yulu 2210 210 m away · 6 min walk Range on this bike 11.2 km Your trip needs 9.3 km ⚠ Tighter than usual This bike clears your trip, but with less spare range than most. If you're planning stops along the way, pick another. Reserve and walk over See other bikes

The one place it speaks

Marginal, not failing. The notice informs but never blocks, and colour is backed by an icon and a bold heading.

ADMITTING FAULT
Ride complete 27 min · Yulu 4471 Riding (27 min) ₹54.00 Adjustment −₹18.00 Total ₹36.00 Adjusted — ₹18 refunded The lock didn't register when you finished at 2:14 pm, so we've removed the extra 9 minutes from your fare. ☆ ☆ ☆ ☆ ☆ Done

The trust moment

The system caught its own failure and fixed the fare before Ananya noticed — stated as fact, with the reason, no apology.

This is Ananya's original bad ride, rebuilt. The same hardware still fails — the cost of it just moved from her to the system.

Still open

Designing with AI is mostly deciding what not to hand over.

01

The restraint is the skill.

Being able to say why a capability isn't AI is what separates a product from a demo.

02

A value that isn't in the prompt isn't in the product.

The asymmetry rule sat in my spec with zero effect until it was written into the prompt.

03

The dangerous bugs live between the layers.

Wireframe, spec and code were each correct — and the product still had a hole in it.

04

Use AI hardest where it can argue with you.

The most valuable model output was "you would not draw retention conclusions from this data" — it cost me the claim I most wanted to make.

05

Failure states are the product.

Trust is built or destroyed in five of seventeen screens.

06

Invisible systems still have to be inspectable.

If the rider never sees the intelligence, someone else has to.

NEXT CASE STUDY 01 — GENIAUS

Turning AI into a co-pilot auditors actually trust.

Read case study →