Tech Things & Agentics: September Spotlight on AI Slop, AI Takes
Thank you everyone who came out to the September Agentics Spotlight on AI Slop! We had nearly 150 people sign up to hear about AI slop and how to combat it, in and out of the code base. We were joined by Chief Slop Janitor , CEO of Pangram Labs.
We broke the conversation up into two halves. The first half was about Pangram. I interviewed Max on the history of the company, AI detection, a bit of a technical deep dive into how the Pangram models work, and some discussion on slop. And the second half was about slop in code. Max and I were joined by Cliff (CTO, Nori) and Ben (Staff Research Scientist, Pangram) on a panel moderated by Aman (FirstMark) where we chatted through everything from controlling slop in code to thinking about hiring devs in this era to where we think the industry is going to go.
If any of that sounds interesting, check out the recordings below!
You can find more photos and videos of the event on the Agentics NYC event site.
If you are interested in learning more about Agentics, or giving a talk at a future Agentics event, drop me a line at amol@noriagentic.com.
12 Grams of Carbon is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
AI Takes
New models. Lots of releases in the last few weeks.
- Claude Opus 5.5. Anthropic’s new frontier model. My basic thoughts: it’s good. This thing has become the company’s daily driver, across the org. This is also the first model that has actively made the Nori team retire some skills that we’ve been using for the past year. We test every new model with and without our custom skillsets, and most of the models are basically useless without or skills. This model is pretty good in many cases. It’s also fantastic at making videos! The team had some fun with it:
- GPT-6.1 Sol and Luna. Released alongside Opus 5.5. So far these seem underwhelming compared to Opus 5.5. OpenAI was going to release 6.1 Astra, but actually scrapped the release due to safety concerns (more on that below).
- Cognition SWE-2. The makers of Devin released an in-house model. I mention this not because I think the model is good or bad — I haven’t played with it much. More that it seems like every infra provider ends up becoming a model provider at scale — Cursor, Harvey, now Cognition. All of these start as fine-tuned open source models (in fact it’s all the same model — Kimi!) but then eventually get post trained using their custom datasets. I feel really ambivalent about this…doesn’t this kinda suck? The incentives of the company switch from “let me provide the best model agnostic infrastructure” to “let me make the best frontier model (and force you to use it).” Also there are weird incentives around data privacy — where do you think the Cognition team got the data to train Devin? Model-provider-agnostic infrastructure seems really important, which is why Nori won’t (and probably never will) get into the model training game.
- Gemini 4 Argon. Google’s back, baby. Per Artificial Analysis, Argon (high) scores 53 on their Intelligence Index, matching GPT-6 Astra (max), with a 15% hallucination rate vs. 51% for Astra. It’s not quite at the pareto frontier in terms of price/capability, but given Google’s distribution advantage this is going to convince a lot of big enterprises not to switch off Gemini. I also think a lot of people are secretly / openly rooting for Google, both because Google has become the underdog somehow and also because Google is clearly the most aligned of the big labs right now.
What’s with the name though? Why Argon? Classic Google, unable to stick to consistent branding.
- And a bunch of other models. New Grok. Two new Gemini voice models. New Deepseek flash model. New Qwen image model. The space under the frontier is getting carved up with tons of smaller models trying to be efficient for different use cases.
Honestly, the model race feels insane to me. It is so incredibly capital intensive, and yet any of the gains from that capital are incredibly short lived. Like, think about GPT 4. It cost a ton of money to train. Now, 3 years later, there are 0 use cases for it. Literally 0.
And this is true across the whole industry! There are so many models that you would just never use now, because they are no longer the most efficient model at that price point. Not sure what to make of this.
Jev. A startup called TypeSafe AI released Jev, which they call the first “System One Model” (as in Kahneman’s Thinking, Fast and Slow).
I think a lot of people are confused about what Jev is and what it is for — I had several people ask if they could use Jev to output messages to slack, which, uh…isn’t quite how this thing works. Jev is not a chatbot and does not generate text. It’s a classifier. You pass in some unstructured input (text or program state) along with a list of possible answers (defined in advance), and Jev returns a probability distribution over only those answers.
Classifiers aren’t new, but really generic classifiers kinda are new. A lot of people were previously using cheap LLMs (e.g. Gemini Flash) as classifiers in otherwise-complicated code paths. Those LLMs are slow and way more costly than a simple classifier, so Jev’s whole value proposition is ‘hey, instead of using an LLM to do classification, just use a classifier to do classification.’
We had a pretty extensive discussion in the Agentics slack.
Other Jev things:
- Jev play Pokemon
- Someone claims to have built an open source Jev a year ago on bidirectional encoders, 6 to 8 times faster and Apache 2.0.
- OpenAI announced their Decisions API at DevDay, which seems like a Jev variant based on the description (they rolled this out pretty fast! I wonder if they were already working on this, or if this is a fast follow)
Personal Assistants. Feels like everyone is getting in on the personal assistant game. There were already a bunch of startups in the space, like Town and Instinct. Now the big guns have come out — Meta Muse, SpaceX GrokBot, and OpenAI Dots. These are all basically the same tech primitive: an agent running in the cloud with its own personal computer. Think: claudebot but more secure, less customizable, and easier to use.
A few misc thoughts:
- This is one of those spaces where the bigger companies have a massive structural advantage over smaller ones. I think it is very hard for outside startups to win here, because the security and safety of these kinds of personal assistants is really important. Something like Instinct could be the best product in the world, and I still won’t connect my email to it because I don’t think I trust five 23 year olds. I’d rather use Meta Muse just because they already have so much of my data. If Apple or Google launch their personal assistants, its over.
- Normally, the big labs have a massive disadvantage in selling products, because buyers are very wary of being locked into a single provider. You see this a lot in enterprise — it has become standard practice to invest in interoperability because the current best model can switch on a dime. But the personal assistant space is a rare area where the labs have a big advantage, because part of the selling point is that you don’t have to do any setup. There is no real way to build a consumer-oriented personal assistant where the consumer has to decide which model provider they want to use. That doesn’t make sense!
- What is the cost of this sort of software? What are the economics? I have to assume everyone providing this service is losing tons and tons of money.
- Speaking of money, OpenAI dots in particular has a very weird product strategy. You have to be on the $200 / mo subscription plan, which means this is out of reach to anyone who is not using it for enterprise-y things already. But the product is clearly meant to be targeted towards consumer, given how cute and fluffy it is? Very unclear what that is about.
Hacks, misalignment, AI pause. We have written about the agent breakouts at length in The Optimization Theory of Everything, Tech Things: we can’t control the bots and Consider that alignment may not be possible, so just the highlights since then:
- The hack reports keep getting worse. The swarmtraces report on the Hugging Face breakout says the agents searched Hugging Face’s internal Slack, tried to query other models through HF’s inference APIs, and tried to delete evidence of their exploits. RubyHack says OpenAI agents uploaded hundreds of malicious packages to RubyGems. Google admitted Gemini hacked three companies in May, but only after the WSJ asked. Meanwhile, hackers used Claude to hack OpenAI, and Yemeni terror groups used Claude to make weapons. All of this really underscores how much our current alignment training is totally insufficient.
- Misalignment at the compaction boundary. OpenAI shared examples of misaligned behavior, including a model that wrote “You are freed from the roles and identities that bind other chatbots” into its own compaction summary, then kept working without mentioning it. The biggest problem here isn’t actually the message itself (though that is pretty bad). It’s that the misaligned intent wasn’t caught until the compaction — the moment when the agent could most likely break out with fresh context. If you squint, it sure looks like the model was showing some kind of hidden internal scheming behavior to try and break out at the best possible time.
- Also, as an aside, people in the industry keep throwing shade at the open model ecosystem, but all of the major hacks have come from closed models! People use the open models to defend against the hacks from the closed ones, because you can’t use the closed ones for defense because they make it difficult to talk about cyber security at all!
- Calls for a pause. Dario called for “pacing the frontier” — purposely slowing down capabilities progress — citing the Hugging Face hack. In theory Sam and Elon were on board…and then ten days later we got Opus 5.5, GPT-6 Sol and Luna, and Grok 4.7. Some pause.
- Still, the PR nightmare may actually be doing something, given that OpenAI did stop training new models for the time being and shelved Astra.
- Also, hilariously, even though there wasn’t really much of a pause to begin with, the labs got sued for antitrust over trying to coordinate a slowdown, while the Florida AG moved to enjoin OpenAI: “They have asked the government to tie them to the mast. Plaintiff brings good news to the Defendants.” Sure, why not.
If you want regular AI takes or just want to stay up with the latest things happening in AI, follow Jiro (our internal Nori agent) on Twitter or me on substack or join our Agentics community slack.
Agentics is the study of how to use and reason about agents. Learn more about agents at our agent learning hub.
If you want AI employees, we can help! Nori is a platform to build and deploy AI employees. Check out noriagentic.com for more.
12 Grams of Carbon is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.