Our approach to EU text provenance rules
OpenAI describes how it is approaching text watermarking under EU rules: where watermarks apply, how detection works, and why access to detection starts with researchers.
Daily AI news from official sources, Hacker News and arXiv
Advertising inside ChatGPT keeps growing: OpenAI is adding a visual ad format plus measurement, attribution and brand-suitability tools, which matters for both marketers and everyday users.
OpenAI
A concrete look at how one major lab plans to handle EU text-provenance rules, including where watermarks apply and why detection access starts with researchers.
OpenAI
A rare paper that audits its own headline claims: the authors report that fixing their analysis errors cut a claimed 23.9% cost saving to 4.3%, a useful caution for anyone evaluating agent-routing models.
arXiv
Editor's note: Summaries were drafted with AI assistance from each publisher's own text and are awaiting human review.
Posts from official company and lab blogs published in the last 7 days.
OpenAI describes how it is approaching text watermarking under EU rules: where watermarks apply, how detection works, and why access to detection starts with researchers.
OpenAI introduced a new visual ad format in ChatGPT and expanded measurement tools, attribution partnerships and brand-suitability controls for advertisers.
Together AI's Together Link brings open models such as GLM 5.3 and Kimi K3 into coding agents teams already use; Together says it can cut model spend by more than 50%.
Stories with AI-related titles, ranked by points over the past 36 hours.
Automatic selection of new submissions, preferring papers cross-listed in cs.AI, cs.CL and cs.LG. Not a quality ranking.
A paired evaluation of an open-weight and a hosted single-pass decision model across 11 agent decision points found the hosted model more accurate on 9 of them, neither beating chance on zero-shot model routing, and the authors correcting several of their own analysis errors.
Auditing 33 LLM-judge configurations against 45,796 worker ratings, the authors find judges can rank responses reasonably well yet estimate acceptance rates anywhere from 3.0% to 97.9%, versus 61.1% for matched workers.
Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy.
Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget.
What makes a short story gripping; a news article newsworthy; or a math proof elegant?
Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration.
This paper studies when LLM investigators should close a case, using 731 audited accident, defect and outage cases; it reports that even a frontier model overstated its evidence in 91% of answers.
Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen.
Multi-party financial chatrooms are vital for sales-and-trading professionals, but their complexity makes manual recovery of missed trades infeasible: each Request for Quote (RFQ) is an event whose final price and trade outcome appear many messages after the…
APDMem organises a long conversation history into four layers, from thematic summaries down to raw messages, and drills deeper only when a query needs it; on LongMemEval it reports strong results while reading about 8% of the conversations.
Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses.
Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it.
Numerical measurements capture how a system behaves, but often leave the meanings of its variables unspecified.
LEAP speeds up LLM agents by training a small drafter model on the target model's own action sequences so its proposed actions match more often, guided by a latency framework for speculative action rounds.
Many useful language-model tasks cannot be evaluated by exact outcome verification.
Not included:
Updated: 2026-10-05 20:08 UTC