why is there a pattern like this

I have been using codex for like 7 months now ever since gpt 5. and ive been using it for my side project ever since one thing i can note is there is always a pattern: gpt 5 dominates my workflow with little hand holding. then gpt 5.1 comes out gpt 5 starts to hallucinate then i go to gpt 5.1 because gpt 5 can no longer do the responsibilities im putting on it. then its the same loop every "frontier release" the model i use can no longer keep up so i have to move to the "latest frontier". gpt 5.6 sol was the sweetspot for me. ive tested stuffs on it from engineered prompts to simple "can you do this and that"(basically acting like the client) and it was doing good. then astra came out and it cant no longer but now astra has a usage problem which made me go back to gpt 5.6 sol and TO MY SURPRISE(not rlly) it can no longer perform the way it was performing before. it requires alot of hand holding now no matter what reasoning i set it to. and the worst part? its using more tokens more than when it was a frontier model if anyone wondering the projects are just minecraft modding but ofc the coding part does not use general coding cuz it has to hook to alot of API like forge or fabric and also mixins into minecraft code itself to make these mods (i dont see degration on general coding tho like sites) but basically this is the best benchmark for me as it does require them to read api (i made a skill file for these so dont worry its not reading the entire library. it has guidlines) but sadly only the frontier can follow instruction and follow the custom skills guideline. if anyone wondering "what changed? how do you know it no longer can do the responsibilities your applying?" at first gpt 5.6 sol can follow the skill i made to a tea. following guidelines using the scripts and mcps to get the right context it needs... but now its disregarding the skillfile and references even if i call it every prompt and just searches the web which it never did before. then i added to my agent.md to not use the web search feature when modding. but the outputs starts being low quality. if you say "skill issue" im not having issue with workflow xd im just having issue with degration and not knowing if i send a prompt if its gonna perform like yesterday... i know people use this professionally so having stability is something that would be best for everyone. knowing when a model becomes weaker would be a great help before people start their work

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论