[AINews] TypeSafe/Jev at >$100M ARR, $7.5B valuation 3 weeks after launch
As you can see in the AINews X recap section below, everyone on earth has cloned the Jev API, but only one company can ever create the category. TypeSafe announced their “Series AI” and Sequoia “leaked” that they crossed 100M ARR in their first week.
Although there are cynics and accusations of astroturfing, we hope it is evident that our Jev pod was 100% authentic.
AI News for 10/8/2026-10/9/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can of email frequencies!
AI Twitter Recap
Decision Models Become a Product Category
- The pattern: Several vendors shipped “decision” models on the same day. These return typed answers (probabilities, picks from a list, scores) in a single forward pass instead of free text. Jev is the reference point everyone benchmarks against, and @scaling01 remarked on how fast the format spread.
- OpenAI Decisions API: Three request types: probability that a condition is true, pick from a list, or score against levels. It accepts text and images, runs on GPT-6 Luna, costs $0.10/M input tokens with no output charge, and is “up to 10x faster” by OpenAI’s own figure (@LearnOpenCV).
- Microsoft-Decision-1: Positioned for LLM judges and screening scientific hypotheses. An early evaluator says decision models still struggle on consistency and complex decisions (@omarsar0).
- Perplexity pplx-decider-v1.1-27b: Claims top Decision Bench accuracy at 94.5% across 1,071 cases, at $0.017 per 1K decisions (@perplexitydevs).
- Cloudflare clef: New clef-omni accepts audio, video, image and text. clef-flash is now cheaper than Jev, and clef overall is about 2x faster (@michellechen). Weights are on Hugging Face.
- Liquid d1: Now on Vercel AI Gateway, with vision support for classify, route and score tasks (@vercel_dev).
- Serving and routing: vLLM Semantic Router’s Decision 2.0 answers multiple questions about one input in one pass, with per-option probabilities (@vllm_project). LangSmith uses Jev as a judge that returns separate typed answers for difficulty and correctness on every trace (@hwchase17).
- Train your own: Unsloth released a free notebook that turns Qwen3.5-4B into a decision model on 8GB of VRAM (@UnslothAI). A walkthrough on Qwen3.5-0.8B reports accuracy rising from 37% to 65% in 60 steps, about 10 minutes on 4GB (@akshay_pachaar).
- Why harnesses want this: Many agent steps are yes/no calls rather than generation. LangChain says routing each task to the cheapest adequate model cut median Open SWE cost per task by 64% (@hwchase17).
- Related research: Apple/CMU’s Selection-based Structured Reasoning (SSR) applies the same idea inside agents (@ZhihuFrontier).
- Method: Six natural-language strategies are scored by length-normalized log-likelihood in one batched forward pass that shares the KV cache.
- Results: Per-turn reasoning latency falls by more than 90%, but end-to-end latency per question falls only 28–54%. On Qwen3-VL-4B with GRPO, average success is 61.37% versus 61.25% for a TAPO+GSPO baseline.
Multi-Agent Orchestration and Coding Tools
- Claude Managed Agents dynamic workflows (public beta): A lead agent writes a phased plan, fans it out to up to 1,000 agents per run, then merges the results. It is enabled with
multiagent_20261001(@ClaudeDevs, config).- Cost warning: Anthropic advises starting with scoped tasks because token use can be high (guidance).
- Claude Code Projects: All waitlisted Pro and Max users were admitted. Each project runs tasks as parallel threads (@ClaudeDevs), and sessions can now run locally (@gem_ray).
- Opus 5.5 fast mode: It has rolled out, but it bills against usage credits and is not included in subscriptions (@theo).
- Do agent teams pay off?: Vals AI ran GPT-6 Sol and Opus 5.5 on Vibe Code Bench, alone and as teams (@ValsAI).
- Results: Teams cost 1.8–5.1x more. Only Sol at medium effort improved significantly, by 7.3 points.
- Behavior: Sol delegated in parallel along architectural lines. Opus ran sequential waves, reaching about 6.8 subagents and roughly 1,140 subagent tool calls per app at max effort, with no significant gain (details).
- Prime Agent rewrites itself in Rust: Over two weeks, a swarm of more than 2,000 agents used 10K+ sandboxes and 200B+ GLM-5.3 tokens. The result reaches usable input about 13x faster and uses 83% less startup memory (@PrimeIntellect). An accompanying essay argues that context limits lead inevitably to swarms (essay).
- Codex updates:
- Windows sandbox: A new mode built on Microsoft Execution Containers (MXC) gives faster setup, network enforcement and granular file controls (@OpenAIDevs).
- Composer predictions: Codex now suggests your next message, in beta for Pro users only (announcement). Some users criticize the Pro-only gating (@Angaisb_).
- Reliability: There were complaints of daylong outages (@dzhng).
- Sentiment: DHH says GPT-6.1 Sol made Codex his primary tool over Claude (@dhh).
- Devin and Grok Bot: Devins can now spawn trees of managed Devins, so wall time tracks the slowest branch rather than the sum (@devindevelopers). Devin also accepts personal ChatGPT plans for GPT usage (@cognition). Separately, Grok Bot gets its own email address for sign-ups and scheduling (@bot).
Model Releases and Independent Evals
- Qwen-Image-2.1-Turbo (open weights): An accelerated checkpoint of the 7B Qwen-Image-2.1. It does 8-step 2K generation and natural-language editing, loads through Diffusers
QwenImage21Pipeline, and launches alongside Pro and Turbo APIs (@Alibaba_Qwen). - StepFun Step 5 Preview: A 600B-total, 27B-active sparse MoE with 1M context and vision (@omarsar0).
- Results: It scores 33.89 on the Hermes Index, matching GPT-6 Luna, and is free on Nous Portal for a week (@NousResearch).
- Availability: It reached #1 on OpenRouter Trending, which measures usage, not quality (@kimmonismus). Open weights are due October 15. Max output was corrected to 64K tokens (correction).
- Upstage Solar Mini 4: A 35B MoE with 3B active, 524K context and 208 tok/s. Its AAII score of 24 is the best at 3B active, within a point of Nemotron 3 Ultra. It is free in Cline (@cline).
- Gemini 4 Argon: Reported at 77.9% on DeepSWE v1.1 versus Opus 5.5’s 74.2%. It ships first to 650+ Fairwind Program defenders at $2/$10 per M tokens (@dl_weekly).
- Signals: Reasoning-effort selectors have appeared in Antigravity (@testingcatalog), and Logan Kilpatrick says “Argon is coming” (@OfficialLoganK).
- Unconfirmed: Business Insider reports that an internal “Carbon” checkpoint approaches Opus 5.5 on coding.
- Speech models: HeyGen Voice tops the Artificial Analysis Controlled Voice TTS arena with an Elo of 1,201, at $30/1M characters and 40 chars/s (@ArtificialAnlys). Whistle is a 16.9MB on-device STT model said to rival Whisper base (@victormustar).
- Multi-turn image editing: Artificial Analysis chained 30 consecutive edits (@ArtificialAnlys).
- Results: Ideogram 4.5 and FLUX 3 edit locally, leaving 95%+ of the image untouched on small edits. GPT Image 2.5 Sunburst re-renders most of the frame each turn, keeping only about 20% unchanged, so it drifts. Nano Banana 2.1 gradually darkens.
- OCR benchmarks: Roboflow’s new benchmark covers 48 models, with GPT-6 Astra leading text localization (@skalskip92). Datalab’s OmniParseBench has 16K tests across 90 languages, and its own model does not rank first (@VikParuchuri).
- Arena roundup: Claude Haiku 5.5 ranks #30 on WebDev at $0.10/$0.50, matching GPT-6 Luna’s price while scoring 6 points higher. Mistral Large 4 sits at #43 on Agent Arena (@arena). On ARC-AGI-3, a new high score of 59.17% (@arcprize).
Research, Training and Inference Systems
- vLLM and SGLang on Vera Rubin: vLLM reports more than 7.8x GB200 throughput on MiniMax M3 at matched interactivity on AgentX. These are early results (@vllm_project).
- Technique: Locality-aware MoE uses CUDA 13.4 locality domains so each SM reads only local HBM, worth up to 1.2x faster MoE decode (details).
- SGLang: Up to 20% faster FP8 MLA at 128K context, and a 5.9% end-to-end gain from MoE tail fusion that removes 276 launches per decode step (@sgl_project).
- SemiAnalysis claims: A preview InferenceX submission shows 3.2x profit per gigawatt and up to 10x performance per dollar versus GB300 (@SemiAnalysis_).
- TRL v1.15: The fused LM head is now on by default and avoids materializing the full logits tensor (@LysandreJik).
- Results: On Gemma 3 1B, GRPO sequence length rises from 28K to 114K and DPO from 10K to 59K. Peak memory at 8K falls 52–82%, and training is up to about 11% faster.
- Data and post-training services:
- Datology Curation Studio: Claims a 6x compute multiplier on 39 open datasets for a 30B MoE (@pratyushmaini). It also cites Thomson-1, trained for $450K, beating GPT-5.6 Sol head-to-head (@arimorcos).
- Tinker: Price cuts of up to 70%, long-context priced the same as short, and GLM-5.3-Flash and DeepSeek-v4.1-Flash added (@tinkerapi).
- DeepSeek periodic weak spots: ByteDance Seed finds that retrieval depends on where a token lands relative to the compression stride (@ZhihuFrontier).
- Evidence: The pattern persists without RoPE or learned gates, and tracks stride length.
- Interpretation: V4.1’s stride of 2 reduces but does not eliminate the effect.
- Agent research:
- Agent plasticity (Meta): Measures held-out gain per learning dollar. The best performers are not the most efficient learners (@omarsar0).
- MIMESIS: A 9B user simulator that beats Opus 5 on behavioral fidelity by 13.4 points (@dair_ai).
- Base-model selection (NVIDIA): Ranks checkpoints by whether the base model can reproduce the “decisive edit,” a signal that tracks post-trained SWE-bench Verified scores (@dair_ai).
评论
?
参与讨论