A 24KB stand-alone HTML-LLM that can generate consistent stories
preview.redd.it/7sknv6omnguh1.png Want to try? Just go here and click "New seed" a lot. You will get different stories, it just requires some clicking - lucky RNG draws: output.jsbin.com/nikupemuta/1 You should even get above 60 tokens per second with a smartphone. FAQ: Why?! Because it's possible, and fun. My original idea was to simply mash MacroStory and LittleBit together, but that did not work at all. Is it relevant? Not a tiny bit ! Even a 3 MB HTML file with external dependencies would be loaded in a second and could serve a way more capable model with a lot less work. Squeezing the model itself to 20 KB at 0.01 KLD was trivial. Saving another 5 KB while maintaining output quality required a lot of time. How? Local Qwen3.8, some ideas, lots of patience. Adaptive quantization and QAT made it happen, LittleBit, Rotation, etc just made it worse. HTML/JS packing with minify, zopfli, and a bunch of structural script changes helped with the result. Was this all made by you? Not at all. This would not have been possible without the 81 KB (FP32) MacroStory model. I "just" tinkered around with it somewhat. Details for those who're interested: Size-reduction rules are very different for a tiny model than for a large one. The embeddings are 60% of the model size, while the tensors are the largest part in a normal-sized LLM. The tensors of tiny models are extremely sensitive to quantization, especially with the Ouro looping transformer format of this LLM. The embeddings had some room though. It's mathematically impossible to save space with 1 bit quants or less with the LittleBit approach here, as the model matrices are just too small for that - the gains would get eaten up by the overhead of the correction data. GGUF K quants would also not help for the same reason: the model matrices are simply too small for the introduced overhead. Even if some tensor size could be reduced by LittleBit: the 3 KB of quantization gain would be eaten up by 3 KB of newly required safetensors metadata. Nobody thinks about metadata sizes in normal-sized LLMs. Yet aside from that: A modest 4 bit LittleBit quant would still break the model. The currently chosen approach meanwhile comes with a convenient 0.04 KLD. This already causes a very occasional duplicate sentence or non-matching story start. The MacroStory model has a handful of dead tokens that could be exploited for further size reduction, as they'd never make it through the sampler on their own. Since you've arrived down here, there's a bonus for you: Local Python inference (YMMV) and the non-packed HTML. Just save this image to disk (important: "Download original image"). I originally wanted to use it here or on imgur, but both wouldn't let me. Open with 7-Zip, WinRAR, etc to unpack it. Or on the console - even on Windows - use either: tar -xf MS256story.png 7z e MS256story.png python -m zipfile -e MS256story.png