[AINews] Pi 1.0, Pi Durable, and AIE NYC

Last call for regular tickets for AI Engineer NYC! See you in 2 weeks!

As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps (for new tickets only, no refunds).


Pi is often mentioned in the same breath as OpenClaw, as we did earlier this year:

but today is time for the increasingly well regarded Earendil, which Pi joined, to have its day in the sun, with both Pi 1.0 and Pi Durable hitting the front page of HN.

Pi 1.0:

Pi Durable ports Pi to TypeScript and externalizes all stateful components of Pi:

  • Crash Survival: Every step is recorded as a checkpointed task. If a process fails or restarts, agents and subagents automatically resume from their last exact state.
  • Portability: It runs anywhere with a JavaScript runtime (like Node, Bun, or Cloudflare) and uses pluggable storage backends (Memory, SQLite, JSONL) and flexible remote or local execution environments.
  • Concurrency: A single harness can run multiple parallel, branching conversations—such as a main channel and separate threads—without blocking one another.
  • Extensibility: Developers can bundle custom system prompts, tools, hooks, and durable tasks (e.g., multi-step checkout processes with rollback capabilities) into installable “Extensions.”
  • Context Management: Automatic background compaction summarizes older messages to maintain token limits without pausing the agent’s active work.
  • Multiplayer & State Sync: Application state (like a to-do list) is stored in documents directly alongside the conversation transcripts, allowing multiple users or UIs to connect, watch, and steer the same agent simultaneously.
  • Hot-Swapping: Tool and extension code can be updated dynamically while the agent is running, with the next tool call automatically picking up the new code.
AI News for 10/01/2026-9/30/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can of email frequencies!

AI Twitter Recap

Frontier and Multimodal Launches: Gemini 4 Argon, GPT-6.1 Sol and FLUX 3

  • Gemini 4 Argon: Google announced a new generation of Gemini, with contributors highlighting revised pretraining mixtures, long-horizon post-training data, and internal applications in memory optimization, code migration and mathematics. These are developer accounts of how the model was built and used—not independent evidence of general superiority (Google researcher).
    • Validation: Google says new Gemini revisions now undergo weeks of testing by thousands of internal software engineers before release (Logan Kilpatrick).
    • Contested readiness: A circulated Bloomberg report attributed coding weaknesses to anonymous insiders; a subsequent post reported a senior DeepMind engineer rejecting that account. Treat the practical coding-quality dispute as unresolved, rather than interpreting either benchmarks or employee reactions as decisive (reported criticism, reported rebuttal).
  • GPT-6.1 Sol: OpenAI’s update is primarily an efficiency story. Sam Altman called it the company’s fastest-growing model and said serving performance had improved after launch-time load problems (update).
    • Measured economics: Artificial Analysis reports $0.72 per Intelligence Index task at maximum effort, versus $1.04 for GPT-6 Sol and $3.26 for Astra. Fewer turns and cheaper cache reads—not simply fewer generated tokens—drive the improvement (results, explanation).
    • Multimodal fix: OpenAI also corrected image encoding for Luna and Sol. Luna gained one Intelligence Index point, including improvements on visual-document and knowledge-work evaluations; Sol changed negligibly (measurement).
  • Solar Mini 4: Upstage’s proprietary text-only reasoning model reports 35B total/3B active parameters, a 1M-token context window and 262K maximum output. Weights are not released, so parameter counts remain vendor-reported (analysis).
    • Pricing: $0.10/$0.40/$0.01 per million input/output/cache-hit tokens.
    • Trade-offs: Artificial Analysis scores it 24 overall and 83% on long-context reasoning, but only 1% on Terminal-Bench 4.0. Despite 208 tokens/s output, approximately 88K output tokens per task produce a 7.1-minute average completion time and roughly five times Luna’s task cost.
  • FLUX 3 Image: Black Forest Labs launched native generation up to 4K, up to ten reference images, bounding-box layout control and targeted multi-turn editing. Preserving every untouched pixel is a vendor capability claim, not independently established here (announcement).
    • Availability: Commercial weights are available; an open-weight variant is promised in coming weeks. Hosted access includes fal and Krea (fal, Krea).
    • Pricing: BFL announced a temporary 50% API discount through October 8, without supplying base prices in these posts (details).
  • Interactive video agents: Tavus introduced Griffin, a video-to-video interaction model. It claims 48% of live participants mistook it for a human, versus under 3% for earlier systems; that result should not be generalized into an unrestricted “Turing test passed” conclusion without the test protocol (announcement).
    • Enterprise deployment: Separately, Synthesia launched Sessions: conversational avatars for roleplay and survey interviews, extending its previous one-way training-video product (launch).
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论