An Update on My World Tap to Talk

Previous updates on this saga can be found here, here, and here.

Blessings continue to abound in the Some Guy household. The biggest being that my son, his mom, his speech therapist, and myself are all exchanging conversational turns several times a day now. Yesterday, I was downstairs with my son pretending to be asleep and in response he would shout “Wake up, Dada!” Shortly after he broke a toy barn and said, “Me in trouble.” We are having meaningful two way conversations every day now that would have been impossible three months ago.

There’s so much capability that seems to be getting unlocked every day that I find myself grinning from ear to ear all the time. It’s a miracle that I don’t let myself dwell on too much, because if I did then I wouldn’t be able to do anything else. I’m holding a thread of goodness and my job is to stay focused on it and keep pulling. Still, my son is finally talking to me and with me! Thank God! Thank Christ! Hallelujah!

He’s been talking at me so much in these last hundred days that I don’t even notice it as remarkable anymore. Though, I am thrilled to be assaulted by continuous requests for pretzels, don’t get me wrong.

If you’re new and haven’t been following: I vibe coded a personalized communication app for my non-verbal autistic child. I did this mostly while driving him to appointments. I’d put Claude in voice mode, spew out a requirements document on the way to the clinic, then dump it into the code interface while getting him out of the van. I gave the device to him and ever since then his speech has done a complete hockey-stick.

I want to be clear, that I’m not saying he suddenly doesn’t have autism anymore. If you were to meet him you would definitely be able to tell. Given where we were 100 days ago, though… this is better than my best reasonable expectation of where he was going to be in two years. Every other thing in my life sort of feels muted in comparison. 100 days ago I had a four and a half year old boy unable to say the word “I” with any consistency, maybe once every three or four months by accident, let alone stomp and holler in the kitchen “I want pretzels! Yes please! I want that please!”

To further caveat this, I suspect he had this ability in him and it was latent. We had been told by his speech therapists that he might develop these abilities between the ages of six and nine. This impression is from numerous conversations with his team where the gist was, “if I had to guess, which I won’t because you’ll put too much hope on it, he’ll be able to communicate with you by this age. But I can’t say that. So instead I’ll maybe shrug and tell you about other cases like his and what happened there, and if you were to guess that is my guess it wouldn’t be bad guess.”

I am now in the process of taking the app I built for my son and turning it into something that any family can set up for a non-verbal loved one.

If you’re looking for productivity tips, I don’t have any other than that 3am exists and you can vibe code while driving. I dropped zero of my other responsibilities to make this happen. Sixty to eighty hours a week at my day job, full parental duties watching the kids after work until bed time, taking people on trips, etc. My only advice is all the normal stuff you hear that no one wants to believe. Make a real effort, keep a good attitude, and be honest if something isn’t working then keep adapting. Never give up. Never surrender.

I did all of this because I love my son and from that perspective it wasn’t very hard to stay motivated.

Without further ado, here is a development timeline for the app.

On May 7th, My wife and I were faced with a choice. Spend $1,700 after insurance to get a dedicated communication device for our son, which we already knew he didn’t like based on how he used it at speech therapy and school. Or, I could take a swing and use the same money to buy an iPad and a MacBook to make him an app myself. I have a good history in AI and Product Management at my day job, so I had a very strong intuition I could introduce a lot of innovations that might help my son in particular.

I had never previously made any kind of app before and had no experience with iOS. I do have enough experience and self-belief to know that I can walk into any kind of mechanical or computer-based job and just figure it out. I don’t attribute this to intelligence but rather a willingness to produce high-volume output that isn’t very good and then throwing it away without getting emotional about it. It’s the “turn on the old faucet to get all the brown water out” theory of adult learning. I very commonly meet people much smarter than myself who produce far less impact because they wait on the sidelines for an opportunity to do a perfect job rather than rolling up their sleeves and contending with the fact that they suck for a few months. The only real advantage I brought into this, other than living in the age of AI coding tools, is that I’ve done the whole “suck” to “the best” journey enough times I don’t burn a lot of calories on anxiety anymore. The best tool you can bring to any task is persistence.

In this case, progress was rapid because I could just ask Claude, “tell me what’s available to use in iOS” and rapid-fire pump for information. If I had been doing this with a team I would have been having to go back and forth to all kinds of people to gather all that knowledge. Being able to ask for all the ramifications and implications in one on-demand place was a huge boon to my productivity.

At this point our son was still in the extremely low utterance phase of language development. He maybe spoke twenty words a day, that may or may not have had anything to do with what was going on around him, and it was more or less impossible to prompt him to speak on purpose. His ability to respond to something in conversation was maybe once every few months and before a short while ago he had only ever spoken to me.

Throughout this story you will find that my wife is both very practical and holds a very high standard for our son, so the outcome of this conversation was: “Prove to me that your idea is better and then we’ll see.”

On May 9th, my wife gave me two hours and I went upstairs to try to figure out how to reach my son. I had been “vibe-coding” other projects with Claude Code for a couple of months by that point, but this mattered a lot and everything else needed to be put on hold. I wasn’t doing something theoretical where I could just sigh and say “Ugh, I need some funding to continue this.” However important I believe my other side-projects to be, I was doing something for one of the most important people in the world to me and I’d be seeing his reaction in real time.

I’d had the idea for a while because what the other devices lacked seemed obvious to me. If you’ve ever been around someone with significant autism then you know their attention is extremely hard to hold. They can also be very literal and small abstractions can put things totally outside of their reach. I wanted to make him a bespoke communication device that looked like his actual world. My son has level three autism and every commercially available device we tried had not worked because he would not pay attention to them. It was hard to get him to pay attention to anything.

We had been preparing for a world where our son might simply never acquire the full ability to communicate, but I had one central observation that wouldn’t go away. Our son could flip through books and watch movies. He had some latent ability to pay attention to his surroundings and even laugh where it was appropriate. His attention often reminded me of a flashlight with the batteries going out. If you jiggle that flashlight or hold it in just the right way, you can still get it to work.

It’s the same way for a lot of kids with similar conditions, and it’s what makes parenting challenging. You can raise your voice or whatever you want, but that doesn’t necessarily mean an autistic kid is even aware you’re speaking. The thing that most consistently got our son’s attention was movies. If I couldn’t move his attention to a communication device, was it possible to move a communication device to where his attention lived?

I went with a very simple three column layout. People on the left. Food, Toys, and Movies in the middle. Verbs on the Right. I generated some images on ChatGPT in his favorite art styles from the books and movies he liked, I recorded a few clips of my voice saying the name of each one, and had everything loaded on an old touchscreen laptop in a single afternoon. I recorded my voice not out of ego, but because I’m the person he pays attention to the most. All of this was a single page of html.

I know this is has a certain Hallmark movie-feeling to it, I know it would seem more gritty and real or whatever if it didn’t work right away. The Hallmark truth is that my son grabbed the laptop from my hands immediately, started staring at each picture in amazement because it looked like a story book of our own lives that came out of one of his movies, and was blown away when he pressed a picture and it spoke with my voice. He had his own father’s voice, right there in his hands!

He pressed a picture of his grandfather’s face over and over again.

And he looked at me with his own voice, and said the longest sentence he had ever said, which was something his grandfather often says to him:

“I really love you a lot.”

I really try hard to not get emotional about all of this in front of him, and I reminded myself that night it was probably echolalia, meaning he didn’t know what it meant. But also, if that’s not a sign from God to keep pulling the thread I don’t know what would be.

So, obviously, I have barely slept in the last three months.

On May 10th, I knew I really had something. So I created an Elevenlabs account. I used this for voice cloning because I couldn’t be bottlenecked by having to create clips of my voice over and over again. I barely had anytime to breathe. In the background of this story, imagine I have the most demanding and time-consuming day job possible. My job is more than full time, taking care of my family is more than full time, and this was a whole third thing. Again, I did almost all the “vibe-coding” for this while driving, because I had a few hours every week where I took my son to appointments that I could turn Claude onto voice mode, spew out a requirements document, and then put it into the coding terminal while I sat in the waiting room. Early mornings and nights I would work on picture generation.

I added additional pictures and a fourth row at the bottom of the board for high frequency words, specifically things like I want, Yes, No, More, Again, Stop, Go, etc. I also figured I didn’t want to do code commits every single time I wanted to push an image so I made a proper back-end on Vercel, which was mostly there already from another project, and gave myself an image upload ability.

If I uploaded an image, then Elevenlabs would take the image name and automatically produce the sound clip. I spent the working session setting this up and adding a bunch more images. I also figured out a way to lock and unlock the board that wasn’t intrusive, so I could do this whenever I wanted without my son getting sidetracked.

That night he asked us for a cheese bagel by pressing the cheese bagel button.

After that my wife and I refused to give him anything unless he asked for it by pressing a picture. We were still on the old touchscreen laptop at this point.

For the record, we had tried flashcards before without having any success.

On May 14th, I added better password locking because it turned out the pop up I made to unlock the board was getting in the way. I didn’t add enough logic for when it should disappear on its own. Now, if I didn’t enter the password quick enough the pop up would go away and return the board to a useable state. This solved less than one fourth of the navigation problems he was having. After the initial success, I wanted to make this a web-first application so that anybody could access it on any device, but in practice this just didn’t work and I spent a big part of this day trying to find ways to make it work that didn’t pan out.

I kept the web app active and live, but only as a back-up in case his device is ever broken. I figured other parents would appreciate the same as otherwise even if it’s not perfect, because if one of these things break you’re just hosed until you get a new device.

There were tons and tons of food requests this day.

On May 17th, I realized I needed to think a lot about what I was doing here and started to research language acquisition. I knew there was a certain sequence in which words were usually gained and I wanted to reproduce it. So I took my first stab at creating something I now called “the Taxonomy.” The Taxonomy is basically a list of words, the age at which they are usually acquired, as well as a prompt that describes how I could make that word or concept relevant to my son in an image. Or at least that’s what it was at this stage. I just copied and pasted prompts into ChatGPT to make images in bulk and then I would do a bulk upload so that I could do stuff like add the numbers one to a hundred and every animal that you see in kids books. This became the single most important data object in the entire project.

I didn’t realize until later it should be called “The Dictionary” and by then I was too committed to the name in all the back-end files. Hey, I literally sometimes fell asleep on my laptop while working on this.

Meanwhile, my son would spend an hour just flipping through every animal and number hitting each one. I added the alphabet as well. We could hold up different puzzles pieces of toys in front of him and he would reliably navigate to the image on his board and click on it. At long last, we were seeing something like an actual foothold! Like I said earlier, the biggest and most immediate change for us was that we knew what he wanted to eat all of the sudden. Which was pretzels, for basically every meal. He gained “language as a menu” ability almost right away.

By this time, we had also gone to his Speech Therapy clinic, everyone had cried to see him using language, and I realized other families were going to want to use it so I started work on a public web page but kept my son’s board as my priority. I started work on a waitlist and figured that I would get back to it later so I locked the whole thing behind a password.

The below image is what I call the taxonomy prompt lab where I would generate images from certain prompts with certain models with certain style reference images. I needed something where I could upload a reference image from anyone and immediately generate a whole custom image set. The taxonomy now also includes things like synonyms and different forms of each word added, so run, runs, running, ran, sprint, sprints, sprinting, etc. I had to get imaginative with things like “Look” to be a picture of the child with binoculars standing in a forestry watch tower, and “climb” being a kid on a mountain with a bunch of gear. I found that Opus and even Fable were both surprisingly bad at doing this automatically. Even after I made a skill! LLM’s just aren’t visual entities, I guess.

On May 20th, I realized that my thinking was too constrained. My son could get sight words but I needed him to move beyond that. We had already had small success with him knowing what a thing was called. The device was speeding it up, but I wanted to pour fuel on this fire. So I made a matching game. He would be presented with three images, hear an utterance, and he’d have to hit the right one. When he didn’t get it right on the last try the correct choice would highlight in yellow and then move forward. I made sure to log the outcomes of all of these games so I could start to examine the data. In a month, his scores more than doubled.

Note: I lost this data across a few merges, because I forgot the number one rule of vibe coding. Which is, “the first thing you need to vibe code is an automatic back-up of your back-end and data.”

I would need more data to know, but either he acquired the words at that time or he figured out how to use the board as a communication device. My best guess is a combination of both. He had some latent ability we unlocked, and also he had never seen an aardvark before.

This went through several iterations, many of which I do not have screenshots of. Initially I thought, “I will always initiate these games from my phone and that way it will be interactive.” But my wife is chasing around our baby all day as well as our son, so in practice this meant that the game only got played when I was done with work. I also realized he was starting to look to the board as not only his voice but his knowledge source for the world. If we gave him something new he would automatically look through the board to see what was there. So I made a button that would just work the way he would expect it to work. In product management, I call this “the button that makes good things happen.” The highest goal of a product manager is to build all buttons so the user experience is “Oh, this is the button that makes good things happens.” Easier to earn clicks that way.

Based on whatever the last category was that he was looking at, the quiz game would start the quiz based on that category. As I said before, if he flunked a certain question before it moved screens it would highlight the correct image in yellow, which is his favorite color, and that way there was a mechanism for improvement.

He seemed to not only intuit it would work this way but expected it to have worked no other way. This happened over multiple days, but still… SUCCESS!

He was no longer just learning words but learning the trick of learning words.

At this point he basically mastered learning sight words. It wasn’t a strange thing anymore. His brain just got it.

On May 21st, I realized I had a problem. I could get my son’s attention on the board. I could even get him to use the board to talk “at me.” But I couldn’t get him to use the board to talk “with me.” I made a web app so I could message his board. If I send a message through that web app, his board would temporarily pause and images from the board would appear in a row as well as the clone of my voice. That way I could message him things like “Are you hungry?” without taking the board away from him, as that just never went well.

I will say I had some success with this, but it was small. At first his mind was blown that it was happening without him having to do it, and I tried getting him to understand that it came from me, but he didn’t pay it much attention. He mostly seemed annoyed it took up his whole screen.

Still, I knew that I wanted to do something the other apps didn’t. I wanted a way to insert other people into his attentional sphere and my brain kept chewing on that problem. I’m never really fully mentally present in my day to day life, but I also never stop thinking about how to solve a problem. Not to spoil things, but it took me a good month of thinking before I figured out a workable solution.

On May 23rd, I wasn’t sure what to build for my son so I started experimenting with nano banana integration to onboard new families and generate new images for my son automatically. I started making myself a lot of background stuff to play with for this because I wanted the user experience to be that you’d point your camera at something and click a button and then it would just magically label the image, put it in the right spot, generate three facts, and have it be completely style consistent with the rest of the board. Again, the design guide is, “oh, this is the button that makes good things happen.”

This is a boring screenshot of what I call a “lab” environment where I went to go play around with what set of prompts and reference images would best accomplish what I wanted. The big challenge here is preventing context bleed which is always imperfect.

This second image represents a failure and a problem I had to really think through. I wanted to default most tiles to control expenses for new parents, but I also wanted it to be the case that if you go through the add-image experience enough times it would offer to re-render other images for you. For instance, if you add “Pizza” with a distinct look it would ask if you want to re-render “P is for Pizza” using that specific pizza. All models have quirks so this took a lot of back and forth to make it work acceptably.

And yes, I have fixed the giant button here.

On May 27th, I built a series of routine messages using the scaffolding I’d built with the web app. I could automatically ask my son at lunch time if he was hungry. I could ask him if he needed to go the bathroom every forty-five minutes. My theory is that if I was prompting him at times he wanted to do something it might get him used to the idea of back and forth language.

But… imagine your father’s disembodied voice emanating from an iPad on the kitchen counter asking if you need to go potty every forty-five minutes.

My wife immediately declared this to be the most annoying thing I have ever done, and since I am aware that being married to me is sort of like being married to a very strange wizard from her perspective I eliminated it from his board but kept the scaffolding.

This was another failed attempt at two way communication.

Era 2: iOS Native

By June 2nd, I was fed up with trying to make this behave as a web app on an old touchscreen laptop. He kept dragging his finger here and there and it would cause things to highlight or close pages and I didn’t have the display controls I needed to lock it down for him. Even when I did my best, he could still accidentally break it with ease. This was one of my biggest complaints about the other apps. It was too easy to access all sorts of things that are of no interest to a kid! A non verbal little kid shouldn’t be able to access a settings menu or be asked to type!

I needed to lock down the whole damn page so he couldn’t leave it on accident. Therefore, we returned to the MacBook and iPad idea. I also made the data I had been tracking available to his therapist and integrated ReSend as my email service so I could send her an invite. That way she could add tiles and also look at his data.

On June 4th, I had Claude Code rework the entire app into a Swift UI, meaning a native iOS application. June 4th was also the first day I learned what “Swift” was, which sent me down another research rabbit hole. I had to buy a MacBook in order to do this as you have to have one in order to make iOS apps using something else called Xcode. Once we did this, I was able to lock the screen using something called accessibility settings and get rid of all the navigation issues. We also got a colorful case for him and he seemed to like this quite a lot. The biggest gain for his learning was that I was then able to put this in his lap while driving and have it quiz him on all kinds of things over and over again, which he enjoyed.

The width of the top bar the background colors etc were all changed later.

On June 7th, I finally made myself a full admin experience. The board was getting unwieldy from all the tiles I kept adding and I knew I needed to set this up so I could manage it for other people. I made myself interactive layers for all the different integrations I was running so I didn’t have to do code commits every time I wanted to add a new voice from eleven labs or update a prompt to gemini or reorder the tiles on the board to be more intuitive. I needed to be able to tinker and experiment.

This is my single biggest piece of advice to anyone who wants to start vibe coding. Build yourself tools to tell how good your work is and to get feedback you can put back into the terminal. Even if you’re building some kind of complex state machine, just have it spit out a report with context that has to do with what the report is so you can just dump it back into Claude code and say “fix this.” It’s like mini Foom!

I made it speak Mandarin for one particular child. Which is another way of saying I will make additional language options available as soon as I have enough revenue to pay a native speaker to validate. In this case the dad accepted the limits. More on that later when we talk about labels.

On June 13th, I knew I wanted a way for my son to compose sentences so I made a drag and click experience that fizzled on reception. He didn’t use it all. I wanted to get to back and forth communication and that was the next thing I had been hoping would help me do it. The hypothesis was “Maybe if he understands what sentences are he will understand me better.” No luck.

He was talking “at me” by that point all the time, but we still couldn’t talk to each other. I wondered if I wasn’t pushing for too much too soon.

Then I dismissed that as a weak thought.

Never give up! Never surrender! Be the voice that cries courage in the darkest night!

I kept trying to solve this problem with no luck. This was the problem I knew I needed to solve if I wanted my son to have anything like a normal life. I knew he had some speech capacity now but I needed to press on this as hard as possible while his brain was still plastic. Every little bit of turn-taking ability I could squeeze from him would mean that much more quality of life later on.

The thing I built, which was a pencil icon that would allow him to tap tiles to string them all together at the top and then press play to make a sentence just fizzled. The idea was okay but the interface was all wrong. I figured some kid out there might connect with it better, so I turned it into an optional setting and turned it off for my son.

On June 30th, I woke up in the middle of the night and slapped my forehead. Of course, I thought, you weren’t imagining big enough! You were thinking way too small! AI is more than image generation and text. AI isn’t just large language models. AI is speech recognition too!

I spent that morning researching how iOS devices natively transform speech into text and figured out how to tap into it from an app. I’d seen this happen before when camping and my brain surfaced the memory. I didn’t want to make assumptions through so I spent some time that morning validating how iOS does speech to text and was pleased to see it was exactly what I needed. All I had to do was grant microphone access to my app. Some genius at apple solved the hardest part of my problem years and years ago.

If I have a weakness, it’s that I don’t let go of problems. If I have a strength, it’s that I don’t let go of problems.

A few hours after I had the idea, I was speaking to my son’s device and my words were appearing in the top bar as the images from his new visual vocabulary. I didn’t have to get out my phone or do anything that broke the parity between his vocabulary and my voice. If I wanted to ask “Apple or banana?” then an apple image, an “or” image, and a banana image would appear. More importantly it was real time. There was no significant lag between my speech and the visualization.

That night we busted through all the way to some new dormant part of his brain. I asked him to choose among several different options for dinner. I made him choose from several different movies. He did it, slowly at first, with confusion, but he did it! Finally, we were two way!

I also discovered I needed several hundred new words of vocabulary.

The first day of this thing unlocked sight words and talking “to me” rather than “with me.” He was able to ask for food, toys, movies and name all kinds of new animals. That alone was huge. Everything I did after that was to accelerate his comprehension and his retention of that kind of information. I made it so he could tap the same images multiple times in a row to learn facts about it and how to use it in speech. But it wasn’t until I made this feature that we started to make actual progress in back and forth. After less than a week of having it, I was able to negotiate with him that he had to eat two waffle fries to have an ice cream cone and he did it!

We left that feature on constantly even though it drained the battery much more quickly. Watching my son react to it was an immense joy. It was like he was learning to hear through his eyes.

I have since made the listening mode much taller.

Era 3: Platform and Tweaking

I want to retract some emotional sleep-deprivation-driven insults against the existing devices. Of course they didn’t do any of this crap, because this crap is hard and they need to make sure their product works for the most common use-cases. There isn’t a big market for “can pay attention to certain forms of media, but doesn’t understand language, and is low but not no verbal.” You can’t pay a bunch of employees to go do that work to have endless customization for everybody. I took a deep breath and reminded myself my user base was probably only a few thousand kids total and no company could put in this much effort to reach a group that small.

So, light a candle to Saint Boris of Cherny for making all of this possible.

I’m going to kiss that man’s bald head one day.

And another big one for Jesus, to whom I pray for guidance on this every night.

And Fable, who is a weird entity that despite not being human or possessing a self seems to have a much better attitude than many people I know, and without whom I could have done none of this.

Around this time I took my son for a walk and he spontaneously started playing “I spy” with me. He must have played it in therapy or at school because I had not ever played it with him before, so I just did my best to play it cool and tried to find things that weren’t green. We’d had false starts before but this was the start of him opening up, and we went from almost accidental conversational turns every few months to every few days, to every day, to multiple times per day.

Again, caveats for other parents who have children with these issues. Some of this was probably latent, and I think he benefited greatly from having a bunch of fuzzy memories from speech therapy. As blessed as I feel, it’s also not like he doesn’t still have challenges.

By this point, I had done everything I thought it was possible for me to reasonably do for my son. The app became a sort of shortcut for his verbal utterances when he couldn’t find the words but I turned on listening mode and kept talking to him over and over again and those are the muscles that he needs to flex and which can’t be automated. Anything else I made in the app would be accelerating his progress, and while I certainly want to do that, this also gave me the mental space to start focusing on other kids who would be using the platform.

On July 1st I started to do the stuff I was the least interested in, which was figuring out how to make this a product you could actually sell and make useful to other people without going broke. The selling part in particular didn’t have the emotional swell I got from making something for my son, and put me in the mental spot of thinking of all the parents I sit in the clinic with and then approaching them to say, “Uh… can I have a couple bucks? API costs aren’t nothing. I know this is super sensitive and you’ll get your hopes up and probably be disappointed.”

For whatever reason, in general —so don’t tell me about one guy you know who doesn’t feel this way— the better you are at building the worse you are at marketing. The part of me that can spew data requirements for two hours straight on a drive is the same part of me that notices slight imperfections in communication that distract from the main point of communication. Which in this case is, “here is what this thing is, here’s what it does, here’s where it may help you, here’s how much this costs” and not at all “here is how this works and here’s my anxiety about if it doesn’t work for you.”

This was a very sobering thought I kept in the back of my head as I started to build out the onboarding for other parents. If I’m going to ask someone to commit money to something that will probably not be more than just some cute pictures for their children, I need to give them the ability to do that cheaply. If they failed out on the other apps what were the odds mine is going to help?

My app also isn’t the only cost a parent is going to have to pay. They’ll need to get an iPad too, and that is going to be somewhere close to $300.00. I don’t want someone to be out $300 dollars before they can even try. Trying should be dirt cheap!

Since Claude Code is a magic wand, I did have it make an android app but alas, I have not yet tested this as I do not have an android device and my wife’s patience with me spending a couple hundred bucks a month on Vercel cpu build time is running dry. I’m going to try to make it available on as many devices as I can as cheaply as I can, including optimizing the devices it can run on, BUT that will take time and you can't let the perfect be the enemy of the good. I did realize my dream for this is something a kid can wear on a lanyard around their neck, on a device slightly larger than a cellphone, with a case that has different sensory toys and fidget spinners attached.

In early July, I spent a good chunk of time thinking through how I wanted the onboarding experience to work. Since it costs me money every time I pull this trigger, I tested it end to end a couple of times but nowhere near the number of times I need to feel comfortable pushing this out at scale. I’ve never had to build a load-balancer before, but I built one for this that I think will work. I just also know I can’t just naively put out an app and expect it to work the same when many different people are trying to generate images at the same time. I need to pressure test.

So I also did stuff that was both sales-y and useful: if listening mode hit a word it didn’t have an image for, it would log the word and see if the taxonomy overall had a word. My expectation is that the taxonomy will eventually become tens of thousands of words long and it’s actually not useful to put all of that on someone’s board if they don’t need it. For an early language learner a lot of that stuff is distraction. But if they use that word constantly having it just sort of appear could be another magical feature and I can definitely do that in a privacy preserving way. I built it this way so I can control my costs and just generate a set of standard styles and parents can pay an ongoing subscription if they want it to be continuously customized.

I also got a speech therapist outside of my son’s clinic to really use it and give me feedback (shout out! you know who you are) and there were some critical things I didn’t have like search. Menus can get long and it can get hard to find words. At first, I was going to build a search bar, but again I stopped myself. Why do the lame version? Well, I did the lame version but I also made it so that if you said the same word multiple times the menus would automatically move and highlight that word.

All the admin and overhead stuff brought home a certain clarifying reality. Building with Fable was almost instant and cost free. The bottleneck was my ideas being coherent enough to be computable and being able to hold a large data object in my head. In fact, if I ever make this a business where I am independently wealthy, I will probably get a starlink terminal on a hiking pack, pop in an AirPod, and then just hike and vibe code because my thinking is always better while walking. None of that made it easier to have to sit down in front of a computer and actually fill out forms and get an LLC and business license.

I did get the app out in front of a few people, which felt good. One person in particular reported using it with their son and when he played the quiz game it became obvious he knew all kinds of words he didn’t say out loud. Out of six people that’s the only positive report I’ve had so far although three people never got back to me and given the difficulty in onboarding I don’t know if they continued. Basically you had to go do a bunch of extra steps to get the actual iOS app and I knew that was costing people.

My current biggest blocker is actually setting up a revolving line of credit so I can accommodate large order surges. It costs me money every time I set one of these up but Apple doesn’t reimburse for 60 days, which means that if a few thousand people download my app all at once I would be broke. A guy named Marshall is supposed to contact me next week and help me set that up. I don’t want someone to see this, get excited, and the next thing I do is turn it off and write an apology post about cash flows.

My own trajectory for this is that I have a much shallower user increase but I want to be prepared for catastrophic success. There aren’t that many kids who need this but I also don’t sit on mid five digits of free cash to carry api costs for image generation if I go viral on autism TikTok.

My wife and I have spent way more money than I’m planning to charge on stuff that has a way lower chance of helping. Part of me prefers slow and steady growth because it’ll be easier to shake out bugs, but who knows what sales velocity will look like.

Era 4 — “Launch polish and the magic”

On August 2nd, One father in China, as I mentioned, reached out eager to get a device that would speak Mandarin. He had reached out before this but we didn’t have a good back and forth to figure things out until later. I deployed my super power of just having a meeting at 3am to get it done. One of my greatest weaknesses is that I can’t say no to a kid. One of my greatest strengths is also that I can’t say no to a kid. So I put that taxonomy doc in Fable and said “translate this to Mandarin please. And then I started figuring out how to get an ElevenLabs voice that speaks Mandarin.

It was easy!

The biggest problem was one I had planned on solving later. I baked the titles of each image into the images themselves but this was stupid for a few reasons. The chief reason this was bad idea is it just gave me one more place for the models to fail and require regeneration, which is an unnecessary expense. They also came out inconsistent. The biggest and dumbest reason for making deterministic labels: it was shockingly easy to make the app speak Chinese. I just needed a good way to make the labels change based on language. If I could do that then I could help kids in almost every language everywhere with almost no other changes!

I was originally not going to have labels at all but that turned around the first time I gave it to his therapist and she didn’t know where to find anything and I thought “oh! Duh! The labels are for the adults not the kids!”

I have tested five languages at this point, which only took one day, but since I don’t speak those languages except Spanish —not well, but at the level of a child who would use my app and also a bunch of curse words I learned when roofing houses— I’m going to need more quality assurance.

Early in August, two kids with autism eloped and died and it cast a pall over our home. My son is also a runner. One boy was eleven and non verbal. He took off from his caretaking facility and his body was found several days later in a pipe. The other was five and snuck out at night and drowned in a pond.

My family comes first, so I paused public facing work and made a feature for my son’s app called Beacon. Basically I established a series of geo fences and connectivity conditions. It’s hard to set up which I don’t care about at this point because it’s for him only. He carries his app around with him like a little suitcase. If he’s not somewhere he’s supposed to be and he has the app with him my phone will freak out and blow up all my alerts to tell me that son is somewhere he’s not supposed to be. Meanwhile his app will intermittently, in a kind voice, tell anyone who might be close that the child holding the device is special needs and his caretaker cannot find him and also to contact the authorities and his parents and it shows our pictures and phone numbers. I’m also testing how good of an idea it is to tether his iPad to my phone so if he passes out of range of me or my wife and he’s not in one of his safe spots it starts escalating from there. There’s similar logic if he’s driving with someone who is not us, so a kidnapper isn’t alerted to the Beacon while the vehicle is in motion,

I am not releasing this at launch because I don’t know how successful the app will be overall and you can’t turn something like that on for people and not have done a huge and expensive amount of quality assurance. You also can’t ever pull it back and say “sorry still need to tinker.” If it ever broke for a reason I could have prevented I would never forgive myself and I would also be sued out of existence. My hope is that if I am successful I can just make this free for everyone with the right disclosures and paperwork. On the positive side if I can save a kid I will do whatever I can to make that happen.

I have some more features roadmapped, like books which you can customize each month so it would be a book of your child doing things like eating foods they don’t currently eat, or going to the dentist, etc. That and Beacon are still only for my child though just because at some point I have to stop developing and put it out so other people can try it out.

What will it Cost and When can I have it?

I’m in the process of figuring all of this out. I’m standing up an early access survey to help me measure demand. You’ll need to be specifically invited for the first hundred or so users. Right now I’m making explainer videos and marketing stuff so I can show what all the price points are and what you get for each one. I’m having my son’s speech team make videos as they’re all very excited about it. Also, I’m sleeping six hours a night now and not having any caffeine. I tried recording an explainer video and I thought “yeah, I need to shift my vibe before I try this again. No one is going to buy something from an exhausted man with a slavic brow line.” This is the part I don’t want to rush. I need to launch in a way that even if I get super busy I can provide meaningful support.

Other devices that do things adjacent to what mine does are in the range of $250 to $300 for a lifetime license, or $9.99 a month. That’s for a standard symbol set, no stylization, no making the tiles look like your child and home, no listening mode, no learning mode, no quiz mode, no voice cloning, no stats tracking, but also a wider overall range of characters than I will support at launch. Like I said, I think I just have a fundamentally different type of device for a fundamentally different person.

I should be significantly cheaper and my aspirational goal is to have a very basic customized communication board at a low price.

If your child can abstract enough to use symbols to get meaningful value from these, I would suggest trying one of the other devices. I have no evidence to support this, but if your child has an “abstract this” muscle you can teach them to flex, and you can get them to pay attention to the device, that seems better.

Still, very exciting and at the end of the day no matter what happens: my son is talking now and that means I won the most meaningful thing of all.

More to come. Don’t let me post again before I have the survey ready.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论