Signal SentryUpdated Oct 2, 16:30 UTCPM Drop
Models · open + commercial

New models, and whether they matter.

Every model story the Drops covered, newest first, each with its sources.

From the Drops

Models

Cloudflare open-sources Clef, a pair of fast decision models for agents

Cloudflare released Clef and Clef-flash on Workers AI and Hugging Face under Apache 2.0. It says both are Jev-API compatible, add vision and a 64k context, and beat Jev on latency. A fine-tuning service is coming.

Why it matters: Decision models give agents cheap, typed choices with probabilities. An open, drop-in alternative to Jev lets builders try the pattern without lock-in.

From Drop #002 · Fri, Oct 2 · Evening
Models

Gemini 4 Argon's $2/$10 price is introductory; Google says it will double

CNBC reports Argon's intro price of $2/$10 per million tokens matches GPT-6.1 Sol, but Google says rates will rise to $4/$20. Access is still limited to select cyber defenders and cloud customers, with no public date.

Why it matters: If you are budgeting for Argon, plan for the doubled rate. Broad access, and with it independent testing, still has no date.

From Drop #002 · Fri, Oct 2 · Evening
Models

GPT-6 Astra Ultrafast is in OpenAI's API, up to 8x faster than Standard, NVIDIA says

NVIDIA says GPT-6 Astra Ultrafast runs on Blackwell GPUs and is available in the OpenAI API and to eligible ChatGPT Work and Codex users, with up to 8x faster token generation than Astra's Standard mode.

Why it matters: Faster generation shortens agent loops of editing, testing and tool calls. Check OpenAI's Ultrafast guide for pricing before switching.

From Drop #002 · Fri, Oct 2 · Evening
Models

AWS open-sources Strands Decider 2B, its own take on Jev-style decision models

TechCrunch reports AWS released Strands Decider 2B, built on Qwen3.5-2B, which picks among preset options and returns a confidence score. It is open source, available now and small enough to run locally.

Why it matters: With AWS and Cloudflare both shipping, small decision models are becoming a standard workflow step to cut LLM cost and latency.

From Drop #002 · Fri, Oct 2 · Evening
Models

NVIDIA's DGX Spark gets a 64GB version from $4,999 on Oct. 23

NVIDIA says the 64GB model, sold by Acer, ASUS, Dell, Gigabyte, HP and MSI, runs models up to 100B parameters on device. Two units can cluster to 128GB for models up to 200B parameters.

Why it matters: A cheaper entry point for running agents and open models locally without cloud dependency, with a path to scale to two units.

From Drop #002 · Fri, Oct 2 · Evening
Models

GPT-6.1 Sol is out at $2/$10 per 1M tokens, pitched as near-Astra for less

OpenAI launched GPT-6.1 Sol at DevDay. OpenRouter lists it at $2 in and $10 out per 1M tokens, the same as GPT-6 Sol, with about 1M tokens of context. AWS says it is now on Bedrock.

Why it matters: OpenAI says it matches GPT-6 Astra ($10/$50) on DeepSWE v1.1 at about a fifth of the cost per task, per AWS. If you pay Astra prices, test Sol on your own work.

From Drop #001 · Wed, Sep 30 · Evening
Models

OpenAI scraps GPT-6.1 Astra's release and says its IPO waits on safety

OpenAI told WIRED that GPT-6.1 Astra fell short of its safety bar and won't ship next month. Sam Altman told reporters OpenAI won't go public until it can make confident safety claims.

Why it matters: Plan around the models you can use today. OpenAI says other Astra models will come later but gave no date.

From Drop #001 · Wed, Sep 30 · Evening