Introducing FlyOCR 🪰 - I trained a fly brain to read a PDF

It uses the full MaleCNS v1.0 fruit fly connectcome. The architecture is inspired by doomfly by @nftechie_

The fly splits a pdf image into individual glyphs, maps pixels into receptor activations, runs simplified current-based dynamics across the 166k neurons and 25m edges in the circuit, applies a compact readout model on the downstream spikes, and concatenates everything into the parsed output.

On reading an actual Microsoft 10-k, the fly gets ~86% over the balance sheet heading, but is largely able to read the numeric values correctly. Over 1.7k+ sampled glyphs (chars+digits) it gets 87% accuracy. With enough training it might match some of the latter-generation MNIST models!

Maybe eventually we’ll replace our doc parsing VLMs with flies.

Full video below.
Repo with full code + report: github.com/jerryjliu/fly_ocr

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论