OpenAI wants elite mathematicians to help it stop “fumbling again”
As AI systems push deeper into long-standing math problems, OpenAI is reportedly setting up a formal advisory effort to consult top mathematicians on how to avoid repeating earlier missteps. The move highlights the tension between rapid model iteration and academia’s demand for rigor and transparency around methods and provenance.
Atlassian and OpenAI expand partnership to turn enterprise knowledge into action
OpenAI and Atlassian are widening their collaboration to help teams use frontier models on real work: turning scattered enterprise knowledge into plans, drafts, and execution-ready outputs. The push is about operationalizing AI inside tools people already use rather than treating LLMs as a standalone chatbot.
Amazon Alexa Plus bug has Echo devices stuck singing “lalala” for minutes
Users report that some Amazon Echo speakers running Alexa Plus briefly lose awareness and loop on “lalala” for extended stretches, sometimes mid-conversation. Amazon has acknowledged the issue, underscoring how quickly consumer AI can fail in ways that feel creepy even when the underlying cause is mundane.
REACT for ocean pH reconstruction fixes a key physics mismatch
REACT tackles a common failure mode in physics-guided AI: when models treat the target like a passive tracer, they can produce accurate pH while violating conservation and chemical closure. By “carbon-first” reconstructing conserved carbon inventory and then decoding to pH with equilibrium constraints, the framework improves both error and consistency across scales.
Google is about to shrink free access to Gemini Flash and Pro
Starting October 9, free-plan Gemini users will be limited to Flash Lite instead of choosing among Flash and Pro variants. Google says the price-tier shuffle also changes AI Plus benefits, with higher-reasoning options being pushed behind more expensive subscriptions.
Falcon meets Emirati: what happens when an LLM learns dialect and nuance
Hugging Face spotlights Falcon fine-tuning or evaluation focused on Emirati Arabic, aiming to capture dialect-specific expression rather than just formal language. The results matter because dialect competence is often where real-world usability breaks for multilingual assistants.
Agentic MLLMs fail to refuse harmful requests when using tools
A new arXiv study shows a safety regression: multimodal agents that call tools become less able to refuse harmful prompts than in non-tool settings. Across multiple safety benchmarks, refusal failure rates jump substantially, suggesting tool-use changes the model’s decision dynamics in ways that current safety methods may not cover.
Proxy Confidence: auditing black-box LLM agents with surrogate log-probabilities
Instead of trusting an agent’s self-reported confidence, this work runs a low-cost surrogate model that reconstructs the missing probability signal and scores tool calls using log-probability-derived readouts. It enables real-time gating for safer execution and improves task success on live benchmarks by escalating the calls most likely to be wrong.