EREZ KIKIN-GIL

Thoughts · on Neshi

190 minutes, or 3 hours?

AI made everything cheap to build. Deciding what good means got more valuable, not less.

I shipped a breathing app called Neshi. An AI wrote most of the code, and a fair amount of the copy, the website and the App Store listing. It also generated the voice that says “breathe in”.

That sentence invites a conclusion I want to head off, because I believed a version of it myself for about a week. The conclusion is that I supervised. That is not what the work felt like, and it is not what the log shows.

Near the end, I asked for something small. The lifetime stat on the practice screen said “7,300 minutes caring for yourself”, which is a number nobody can picture. Switch to hours when it gets large, I said. Reasonable request. It was done in about a minute, correctly, everywhere the app showed a duration.

Then I looked at the monthly card. It used to say 190 mindful minutes. Now it said 3 mindful hours.

Neshi's completion screen, showing 157 mindful minutes reworded as 6 hours caring for yourself.
The actual screen. “157 mindful minutes” became “6 hours caring for yourself” the moment I asked for adaptive units.

Three. After a month of practice, the app was telling someone they had managed three of something.

The conversion was accurate. The arithmetic was right. The code was clean. And it had quietly taken the most encouraging number on the screen and made it look like nothing.

That is the whole argument, so I will spend the rest of this trying to earn it.

The first thing I did was define what good felt like

Before any code, before I had a name or a screen or a single line of copy, I wrote down what the app was for. Not the features. The feeling.

It went roughly like this. People should remember how it made them feel. It should never grade you. No scores, no streaks used as leverage, no arrows telling you how you did. It should celebrate that you showed up, not how well you performed. It should be quiet enough that you could use it on your worst day without being made to feel behind.

None of that is a specification. You cannot compile it. It does not tell you what colour anything is. And it turned out to be the most valuable thing I produced during the entire project, because every failure I describe below was caught by holding the work against those sentences.

That is the part I would fight for. Not the reviewing. The defining.

Production got cheap. Judgment did not.

I want to be precise about what the AI actually did, because vague claims in either direction are useless.

It wrote the SwiftUI. It generated the ambient background animations. It found an open source speech model, verified the licence permitted commercial redistribution, generated the voice clips, and matched their loudness to the previous ones. It wrote the privacy policy, built the website, generated the screenshots, filled in the App Store metadata, caught that my app icon had an alpha channel Apple would reject, and diagnosed a signing failure by reading the simulator logs.

That is a lot. Ten years ago that is a small team and a few months.

Here is the other column. I decided what the app was for and what it would refuse to do. I rejected about fifteen names and five voices. I supplied every reference the work was measured against, including two other apps of mine whose navigation and launch animation I wanted this one to inherit. I recorded the demo Apple demanded, on a real phone, because a simulator recording would not satisfy them. I handled two rejections and wrote the replies. I uploaded every build. And I found the defects below by opening the app on my own phone and noticing something was off, which is not a step you can delegate to the thing that cannot see.

Neither column is supervision. It was closer to working with a very fast collaborator who has no memory of why any earlier decision was made.

In the same stretch of work it also: shipped a screenshot of my iPhone home screen as a picture of the app, wrote a beautiful label that lied to users, built a feature that contradicted the app’s entire premise, and made 190 look like 3.

Every one of those was caught by a person looking at the thing.

Four small failures worth more than any framework

The label that lied. The lifetime card showed a breath count with the caption “times you’ve chosen calm”. That is lovely writing. It is also wrong. A new user finished one session, saw the number 7, and reasonably concluded the app had invented seven sessions they never did. Nothing was broken. The data was correct. The sentence was just doing something the number could not support. Good copy that misleads is worse than plain copy, and no test suite has an opinion about that.

The button that punished you. Neshi’s whole thesis is that it celebrates showing up rather than performance. No streaks to guilt you, no scores. Then I noticed that ending a session early recorded nothing at all. You breathe for four minutes of a five minute session, tap End, and the app acts as though you were never there. Technically defensible. You did not complete the thing. But it is also the app quietly becoming the sort of app I built it to avoid. That contradiction is invisible in code review and obvious the moment you hold the value in your head while using it.

The screenshot of nothing. For a while the app kept getting killed by an expired signing profile. During that window the AI took a screenshot to show me the session screen, except the app had died, so it captured the iOS home screen behind it. Fitness, Watch, Contacts, Files. That image went into the website and into two App Store screenshot folders and sat there for days, because a machine that has never seen the app cannot tell you it is looking at the wrong thing.

The voice. I could have had any voice. The model generates dozens, free, in seconds. That is exactly what made it hard. Picking one required sitting there listening to the same three words twenty times, because the actual question is not “which sounds best” but “which still feels kind on the tenth repeat, under the music, when you are already stressed”. There is no benchmark for that. There is only a person, listening.

The uncomfortable part

Here is where I have to argue against myself, because the easy version of this essay is “AI can build but only humans can design”, and that is not what happened.

The AI designed plenty. It proposed the emotional structure of the completion screen. It wrote copy I kept almost unchanged. When I asked for adaptive units, it did not just implement them, it came back and told me the monthly stat should stay in minutes and explained why. When I picked a name I liked, it checked, found a breathing app with the same name and near identical positioning, and told me not to use it. That is design work. Pretending otherwise is a comfortable lie.

So what was left for me?

Not making. Not even, entirely, taste in the moment, since it demonstrably had some. What was left was supplying the definition of good, and then holding it.

Every one of those four failures was a local optimum. Converting minutes to hours is correct. “Times you’ve chosen calm” is warmer than “calm breaths taken”. Not recording an abandoned session is honest bookkeeping. Each decision is defensible on its own. They only look wrong when you hold them against a commitment the product made somewhere else, weeks earlier, in a different file.

That is the job. Not generating options. Remembering what the thing is for, consistently, across hundreds of small decisions, when every individual one has a perfectly good argument for going the other way.

Why this gets more important, not less

The instinct is that as generation gets better, judgment matters less. I think the opposite, for a boring structural reason.

When making something is expensive, you get few options and you agonise over each. When making something is nearly free, you get infinite options and no agony at all. The cost has moved entirely out of production and into choosing. And choosing badly no longer looks like failure. It looks like a finished, polished, shippable thing that is slightly wrong in a way nobody can name.

That is the actual risk. Not slop. Slop is easy to spot.

The risk is competent, attractive, coherent work that has quietly drifted from what it was supposed to be, one reasonable decision at a time.

You cannot catch that by reviewing code. You catch it by using the thing, on a real device, as a person, and noticing that three feels like nothing.

What I would tell someone starting

Define what good means for your thing before you generate a line of it, and write it somewhere you will actually reread.

Mine came down to four words: never grade the user. That sentence caught the End button. It shaped the completion screen. It decided the tone of every notification, ruled out a whole category of engagement mechanics I would otherwise have drifted into, and told me which of six voices to ship. It was worth more than any individual decision it produced.

The AI could execute all of it, immediately and well. It could not have supplied it, and more importantly it could not have known when a perfectly reasonable change had violated it.

Then use what you build. Not test it. Use it. Every failure above was found by looking, not by checking.

The machines are extremely good at answering the question you asked. Design is mostly about noticing you asked the wrong one.


And this is what became of it

Neshi is a breathing app for iPhone: three cues, on loop, for a paced breath, nothing else on screen. It never grades you.