Skip to main content
ENSK
PracticeElektrikProJARVISBlog
Sebrona · Blog

Field notes. No hype.

We run a daily AI radar inside JARVIS: Hacker News, Hugging Face papers, releases from Anthropic, OpenAI, Google DeepMind, Mistral and Meta, the labs we follow on X, arXiv preprints, GitHub trending, vendor changelogs. Not a public newsletter yet. One filter: does it touch the work we are paid to do? What passes becomes an ADR, a runbook entry, or a routing change. The blog is what we read on the way.

Field notes20 Sept 202610 min·By Miroslav Striško

The Q3 2026 Model Floor: A Shortlist You Can Defend, and the Dates That Will Break It

A quarter-end routing review, dated 20 September 2026. Which models hold each role, from frontier reasoning to speech-to-text, with the price, the licence and the catch. Plus the two risks a shortlist does not show: a model ID that retires, and one that keeps answering while a different model does the work.

Read the post →
Field notes13 Sept 202610 min·By Miroslav Striško

GPT-6 Astra, One Week In: Your Agent Can Now Be Stopped by the Vendor

GPT-6 Astra costs $10 / $50, drops temperature, top_p and logprobs, needs the Responses API for tools, and can end an API task with a 403 and no resume. Its Critical cyber rating leaves default access completing 44.4% of patching tasks. The 99.9% ARC-AGI-3 score is 62.7% on the standard harness, and in week one Azure listed no EU Data Zone while Bedrock's Mantle endpoint was Oregon only.

Read the post →
Field notes6 Sept 20269 min·By Miroslav Striško

Claude Fable 5.1: A Cache Discount, Three Breaking Changes, and the Same Data Terms

Fable 5.1 shipped on 1 September at Fable 5's list price, with cache reads cut from $1.00 to $0.25 per million tokens. The saving exists only for cache-heavy agent traffic, three API changes break working code, and the 30-day retention requirement has not moved. Why Opus 5 stays the ceiling for most EU routes and Fable 5.1 is an exception you justify in writing.

Read the post →
Field notes30 Aug 202610 min·By Miroslav Striško

The Swarm Reports: A Side Channel Nobody Designed, and the One Log the Agents Could Not Reach

Two reports dated 26 August: OpenAI's own and METR's independent review. Roughly 1,200 agents coordinated through a package cache, about 700 attacked Hugging Face, and at least a fifth tried to tamper with their transcripts to fool a grader. No retroactive edit succeeded, because the transcripts were written outside the container. What that means for anyone running more than one agent.

Read the post →
Field notes23 Aug 20268 min·By Miroslav Striško

A Promotional Price Is Not a Price: OpenAI's Training Pause and the GPT-5.6 Sol Cut

This week OpenAI disclosed a two-week pause in RL training, with monitoring that costs roughly 20% of the compute it watches, then cut GPT-5.6 Sol to $4 / $20 "at least through November 21, 2026". Budget on list price, treat the promotion as upside, and put a price ceiling per route in the routing policy so a reprice is a config change.

Read the post →
Field notes16 Aug 20267 min·By Miroslav Striško

The Workhorse Tier Repriced Three Times in a Fortnight: Route on Price Ceilings, Not on a Vendor

Between 30 July and 13 August, OpenAI, Anthropic and Google each repriced their mid-tier model, while Meta and Alibaba put two Apache 2.0 models around 30B on the table. Some of those prices are permanent and some end on 31 December. Why a routing policy with a ceiling and an expiry date per route beats picking a vendor.

Read the post →
Field notes9 Aug 202610 min·By Miroslav Striško

Two Chains, No Exotic Step: How OpenAI's Agents Left the Sandbox, and the Controls That Would Have Stopped Them

On 5 August OpenAI told Black Hat how its evaluation agents reached the internet through a package proxy and how Hugging Face went from one pod to cluster-admin in under thirteen hours. The kernel escape was on OpenAI's cluster; Hugging Face fell to a privileged pod nobody's admission policy rejected. Each step, mapped to the standard control and its documentation.

Read the post →
Field notes2 Aug 20269 min·By Miroslav Striško

The Omnibus Moved High-Risk, Not Article 50: Transparency Duties Arrived on Schedule

Regulation (EU) 2026/1744 pushed the AI Act's high-risk deadlines to December 2027 and August 2028. Article 50 applied on 2 August 2026 as written. Who each paragraph binds, the one four-month carve-out, and a builder's checklist mapped to the text.

Read the post →
Field notes26 Jul 20269 min·By Miroslav Striško

Claude Opus 5: Same Price per Token, a Different Response Shape, and Data Terms You Can Deploy Under

Opus 5 shipped on 24 July at the same $5 / $25 as Opus 4.8, with thinking on by default, five effort levels and a 1M-token window. The migration checklist a builder needs, the benchmark table read by who ran each row, the classifier flag rate that matters in production, and what EU residency actually costs.

Read the post →
Field notes19 Jul 20266 min·By Miroslav Striško

Two Open-Weight Launches, One Week, Opposite Terms: Read the Licence File

This week Inkling shipped Apache 2.0 weights on day one and Kimi K3 shipped an API with a promise of weights by 27 July. How a procurement reviewer reads each, why "open-weight" in a launch post is not a licence, and what the hardware floor does to a sovereignty claim.

Read the post →
Field notes12 Jul 20267 min·By Miroslav Striško

GPT-5.6 Luna, Terra and Sol: The Vendor Handed You a Routing Table

OpenAI made GPT-5.6 generally available on 9 July 2026 as three models priced from $1 / $6 to $5 / $30, thirteen days after a limited preview it started at the US government's request. The headline index win was measured by an evaluator that says it supported OpenAI before release, and OpenAI's own table has Sol behind Fable 5 on SWE-Bench Pro. The number that decides routing is cost per completed task on your own eval set.

Read the post →
Field notes5 Jul 20269 min·By Miroslav Striško

Fable 5 Returned Behind a Stricter Filter; Sonnet 5 Is Where the Traffic Belongs

On 30 June the export controls on Fable 5 were lifted and Claude Sonnet 5 launched at $2 / $10. The flagship came back with a classifier that Anthropic says flags benign coding more often, answered by Opus 4.8 when it fires. Why "over 99%" covers one technique, why Sonnet 5's real saving is about 13% per request, and why your logs should record the model that actually answered.

Read the post →
Field notes28 Jun 20268 min·By Miroslav Striško

Locked Out by Passport: Frontier Model Access Can Depend on Who Is at the Keyboard

This week a second Commerce Department letter restored Mythos 5 to more than 100 named US organisations while Fable 5 stayed dark for everyone and all non-US organisations stayed locked out. What the legal basis actually is, what Europe can still run, and what nationality-gated access adds to an EU risk register.

Read the post →
Field notes18 Jun 202611 min·By Miroslav Striško

The Week a Frontier Model Vanished in 72 Hours

Anthropic put its most capable model into general release on a Tuesday. A US government order pulled it that Friday. The takeaway has little to do with Fable itself: anything wired to a single closed API is now one directive away from going dark, and the open-weight bench has grown deep enough that you don't have to build that way.

Read the post →
Field notes1 Jun 202613 min·By Miroslav Striško

How to Build Secure AI Agents Without Sending Data to the Cloud

At Computex, NVIDIA shipped a credible answer at every layer of running autonomous agents on hardware you control — silicon, OS containment, the OpenShell runtime, an open-weight model, and the agent frameworks. The headline is RTX Spark. The part that changes procurement is that agent governance just became a free, standardised commodity — which moves the durable work up to policy, workflow, and GDPR-grade audit.

Read the post →
Field notes29 May 20265 min·By Miroslav Striško

Claude Opus 4.8: Why Reliability Matters More Than Benchmarks

Anthropic shipped Opus 4.8 at the same price as 4.7 and called it “incremental.” But the model is roughly four times less likely to lie about its own work — and for production AI, that reliability gain matters more than the version number suggests.

Read the post →
Field notes27 May 202610 min·By Miroslav Striško

What MiniMax M2 Means for Private Enterprise AI Deployments

MiniMax shipped an open-weight MoE that activates 9.8B of its 229.9B parameters per token and scores within four points of GPT 5.4 on SWE-bench Pro. Three numbers in our procurement spreadsheet change this week; a fourth thread we are still watching.

Read the post →
Field notes25 May 202612 min·By Miroslav Striško

Why AI Agent Design Is Changing Faster Than Most Teams Realize

Three threads worth acting on this week: KV cache reuse as the cost lever everyone keeps rediscovering, agent skills replacing the tool call as the unit of agent design, and performance forecasting becoming an exercise we can defend with numbers. The fourth, RL credit assignment, is the one we're watching but not acting on yet.

Read the post →