AI News Daily Digest (26-08-14)

Previewing Ultrafast: GPT-5.6 Sol runs up to 14x faster in a new OpenAI API tier

OpenAI’s new “Ultrafast” API service tier pushes GPT-5.6 Sol to dramatically higher throughput, with the headline claim of up to 14x speed and as much as 750 output tokens per second. The pitch is simple – make low-latency generation feel responsive enough for interactive apps, not just batch workflows.

Read the full article here

Suno Studio 2.0 adds MIDI – but still no third-party plugins

Suno’s Studio 2.0 takes a big step toward being a real DAW-style workflow by adding MIDI, which should unlock tighter control for editors, producers, and iterative arrangement. The catch – Suno isn’t yet opening the door to VST-style expansion, leaving users with a built-in synth and custom effects instead of a plug-in ecosystem.

Read the full article here

Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts

This paper shows that when MoE routers face quantization noise (like 4-bit KV-cache quantization), token routing decisions can flip across expert boundaries, causing real performance damage. More importantly, it demonstrates a nasty asymmetry – you can often detect that a flip happened, but you can’t reliably predict whether the flip was harmful or helpful, which severely limits “selective repair” strategies.

Read the full article here

What We Learned by Reproducing 2,200 papers from ICML

Hugging Face reports from an ambitious attempt to reproduce 2,200 ICML papers, turning academic “it works on our setup” claims into measurable, dataset-level reality. The key takeaway is less about any single model and more about repeatability – missing details, unstable training setups, and evaluation drift frequently decide whether results actually survive independent runs.

Read the full article here

Flock tightens access rules after a surveillance backlash

Technology Review says Flock is adjusting police access to its nationwide network of license-plate readers in response to growing pressure over surveillance practices. The update is framed as a direct response to contract losses and public criticism, with changes aimed at addressing how the system is used and who can query it.

Read the full article here

Microsoft merges Copilot chat and 365 Copilot into one unified “super app”

Microsoft is consolidating its consumer and commercial Copilot apps into a single interface starting with the Copilot + Microsoft 365 Copilot experience. The change removes duplicated icons and is positioned as one consistent place for chat, image generation, and work-specific features.

Read the full article here

AutoWorldModel-Bench: a benchmark for closed-loop world-model research by coding agents

AutoWorldModel-Bench tests whether frontier coding agents can improve a provided world-model under a fixed compute budget, without being guided by a predetermined “engineering-to-spec” objective. The benchmark found that agents often make non-trivial research-style changes – like new objectives and representations – suggesting a more open-ended way to evaluate agent scientific contribution.

Read the full article here

Does Google even want to win at AI? The Decoder podcast debates DeepMind’s reshuffle

In this episode of Decoder, The Verge digs into Google DeepMind’s leadership shakeup and asks whether the reorganization signals a retreat from frontier ambition. The discussion frames the question as both strategy and culture – what happens when major research talent leaves and productization takes priority over long-horizon breakthroughs.

Read the full article here