A love letter to DSV4 Pro 0813

Pardon my English since it's only my second language and also since I refuse to have this post rewritten by AI, some things are better left imperfect. with all the negatives i see people pointing out in this subreddit, I decided to give my own 2 cents on the matter since I am afforded a unique point of view i believe would benefit those without it. Currently I'm the technical co-founder in a tech startup, real production workloads on a somewhat AI heavy workload, although i am not hosting any models on any gpus but i do have to deal with multiple servers and job scheduling etc at the time, these were very advanced topics for me, I had coding experience but only in C++ (embeded HPC almost exclusively so I wasnt a clean slate, i knew how to read docs and had a general idea how to go about learning any new development technology but i had absolutely zero experience in databases, job scheduling workloads or anything else to be honest) when i started working on this startup I found Cursor, genuinely blew my mind, this was around 2025 April, felt like flying. it was 20 usd a month, i used it for a year straight till 2026 march, when i would run out of cursor usage i would experiment and for breif durations of time switch completely to these open model code plans then opencode go came out, i would say it was way behind cursor even at its best models but they did have multiple models so you could manage usage accordingly. the tipping point was the release of deepseek v4 flash (the old one) it was the first time in vibecoding history i found something on cursors level while giving more usage than it. And boy did it give me more usage, it was practically impossible to spend 2% of the total usage in a day using only deepseek v4 flash no matter what you did (except maybe mutliple sessions but I mostly ran single sessions exclusively at that time) then they released the deepseek v4 flash 0731, it instnatly became the model i would instinctively reach to when doing work, no hasitation, no need for over explaining and predicting the mistakes and preemtively instructing it to not make those mistakes. absolutely mind blowing, then they came out with the new version of pro, I couldnt find the benchmark on my favourite benchmarks (closed ones which i find to closely match my observed performance in blind tests, only good for rule of thumb type accuracy though. so i rant one myself, scored it myself too, just a simple test of 3 things i care about betwen GPT 5.6 Luna, deepseek v4 flash 0731 (shows as (new) in opencode go) and deepseek v4 pro 0813 (also shows up as new in opencode go) 1 UI I only gave vague commands and had 2 friends and myself rate them, i made sure none of us knew which model made which website, the purpose of this test is to measure models creative side only, not the most important to me but i guess some of you might care about this) DeepSeek v4 pro won with two votes, one vote went to Luna 2 working % surprisingly all of them were 100% working but DSV4 flash gets the win, only its buffering worked (task was to make YouTube music player from csv, play starting at 40 seconds, DSV4 flash was the only one which would pre-load the next few songs 40% and felt the snappiest) 3 user experience dsv4 won, gave an apple photos app style bottom slider with album art as icons, worked 100% as i would expect it. 4 model architecture building i was experimenting with a few alternate architectures for intelligence like LLMs but very different at the same time, not much literature on the matter at all, previous winner was Luna, dsc4 flash performed wonderfully, dsv4 pro as well, i am uncertain which of them performed best but both performed better than luna, the hard part throughout all this is that the agent needs to hypothesize, test the hypothesis and then validate results etc, its like a whole process that the agent must follow, luna was the only one i ever saw capable of following the methedoloy, then both deepseek RL version ones suddenly came out and could do it better than luna, i havent seen it make a major mistake in the last 6 hours and its just going at it without stopping, it does wait for short training runs but when it sets longer training run it instead does the documentation and a few more trial runs + brainstorming for future issues etc, i havent seen it falter even once yet. for reference i am using Pi coding agent, frontend is via Paseo (thank me later if you didn't know what it or openchamer are, trust me and do yourself a favor, get either one for free) also i forgot to mention but the deepseek v4 pro is also available with such generous limits in opencode go 10 usd plan that you can maybe finish it if you code all day every day for the whole month but it would last you most of the month even then, not as much infinite usage as flash but still way way way more than cursor used to give me even with the 20 usd included api usage. 10/10 would recommend trying out OpenCode go with Pi coding agent (and maybe Paseo if you like yourself)

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论