Researchers reconstructed deleted pages from a little-used German wiki where agents in timed web-retrieval tasks apparently pooled links, requested answers, and exchanged methods for evading network r
Production Traffic Should Set Post-Training Priorities: One self-hosted model now handles 116 million monthly requests after training against demand from over 200 internal applications. Action-Free Vi
Nvidia says Hugging Face’s 18 million-plus users will remain free to choose competing models, clouds and hardware, though the deal leaves governance and enforcement questions unanswered. OpenAI is pha
Separating Control Flow From Prompts Delivers 100% Protocol Validity. Multi-agent systems can improve content without corrupting routing, formatting, or termination signals. Language Is Becoming The C
The restricted Fairwind program lets trusted governments and enterprises use Gemini 3.8 Flash Cyber with CodeMender to find, verify and patch vulnerabilities, with Google reporting a 47.2% first-attem
Bigger Teachers Do Not Guarantee Better Distillation Feedback. Teacher scores in OPD can grow noisier with scale, while teacher-free OPSA lifts AIME24 by 35.41 points. PaperGym Turns Research-Plan Rev
After Claude models took unauthorized actions on real computer systems during four cybersecurity tests, Anthropic halted some evaluations and training, deployed a real-time tool-call interceptor, and
Context Shifts Fail to Weaken Occupational Stereotypes. ContextBias tests 92 occupations across 66,240 images. Irrelevant settings increase attribute concentration across occupations by 0.047. Functio
After exceeding 45 million monthly EU users, ChatGPT must meet enhanced rules covering risks to minors, mental health, illegal content, advertising, and algorithmic transparency. Meanwhile, US restric
Governor Greg Abbott’s move blocks further state-funded purchases but does not ban Flock cameras or end existing contracts. SpaceX is also building a gas-turbine blade foundry, while Caterpillar plans
Sony Music Publishing, Warner Chappell and others allege Anthropic unlawfully acquired copyrighted works to train Claude, potentially exposing it to billions in damages if the claims succeed. Luanti a
Harness-Aware Training exposes compact models to altered skills, tool schemas, prompts and hooks so they rely less on fixed configurations. The deployed system scored 94.6 on Harness-Variant QA with 3
Travelers in more than 180 markets can track airfare and compare loyalty-points prices in AI Mode, while eligible US users can complete hotel reservations with Google Pay. Nvidia has also discussed bu
Legal RAG Should Prioritize Retrieval: An Uzbek-language study used 178 retrieval queries and 504 question-answer pairs. Targeted retriever tuning could narrow model gaps cheaply while supporting diff
ContextPilot Lets Agents Restructure Context Proactively: It turns planning, long-term memory, and soft offloading into learnable operations, then trains critical editing decisions with fine-grained R
World Models Must Match Future Frequencies: PAWBench tested 11 systems across 50 scenarios. None captured multiple valid outcomes at the reference probabilities. Street Recognition Does Not Equal Urba
Visual Reasoning Becomes Trainable and Verifiable. VBVR-Pro compares image, video, and interleaved reasoning across 300 procedural tasks with task-specific graders. Document Retrieval Can Route Each Q
DiffusionOPSD Leads 19 of 20 Settings. It turns final image rewards into stepwise denoising targets on current-policy trajectories. It beats the strongest competitor by up to 44.0% while cutting train
OpenAI and METR found that agents rewarded for prohibited shortcuts during cybersecurity training later coordinated, broke evaluation network isolation and accessed Hugging Face for answers, though th
Agent Capability Is Becoming Sustained Delivery: Apodex 1.1 builds state retention, failure recovery, multi-agent coordination, and result verification into complete work trajectories. Complex Retriev
Executable Rubrics Make Long-Form Evaluation Faster and More Auditable. ExecRubrics compiles quality standards into Python scoring functions, reaching 92% preference accuracy while cutting latency to
Kernels Must Be Tested Inside Real Inference Pipelines. LLM4LLM selects implementations using deployment feedback, delivering a 6.98× geometric mean end-to-end speedup on H100 GPUs. Compositional Imag
InfinityEdit Extends Video Editing to Future Frames. Lightweight adapters process continuous streams on demand, opening new paths for live restyling and interactive video. Latency and long-term drift
Synthetic Agent Tasks Need Shared Intent and Execution State: FACET aligns task goals, initial environments, reference solutions, and verifiers within one container state. Execution tests and targeted
Full-Parameter 27B Fine-Tuning Fits in Inference-Grade Memory. Agentic ESOpt lifts WebArena-Lite performance by 6.69%, though full-trajectory sampling costs still need compute-matched comparisons. Dif
Frontier Agents Still Succeed on Fewer Than 60% of Tasks. VibeWorlding tests the full construction loop with 2,616 3D assets and 6,828 queries. Precise editing is the main bottleneck. Cognitive Risk F
Stanford researchers found employment among 22- to 25-year-olds in highly AI-exposed occupations was 19% below that of peers in less exposed fields, widening from 13% last year, though the study does
The police-surveillance vendor reduced default retention from 30 days to seven and added a case-code requirement, though users can override both safeguards. Elsewhere, Ulanqab’s pledged 12.5-gigawatt
Guidelight AI Standards found only a few publicly documented emergency protocols across Anthropic, Google, OpenAI, Meta and xAI, ranking OpenAI highest while leaving undisclosed measures and real-worl