A model guide for the GPT-6 family
OpenAI published guidance for startups on choosing among GPT-6 family models, setting reasoning effort, and preparing prompts, tools and workflows for production.
Daily AI news from official sources, Hacker News and arXiv
Unauthorized distillation of model outputs is a growing security concern for AI labs; this is OpenAI's own account of detecting and disrupting one campaign.
OpenAI
An 8B open-weights model built specifically for cited scientific reports that teams can run on their own infrastructure.
Ai2 (Allen Institute for AI)
Cloudflare's post on open-weight decision models and an RL fine-tuning platform was the most-upvoted AI story on Hacker News in this edition's 36-hour window.
Hacker News
Editor's note: Summaries were drafted with AI assistance from each publisher's own text and are awaiting human review.
Posts from official company and lab blogs published in the last 7 days.
OpenAI published guidance for startups on choosing among GPT-6 family models, setting reasoning effort, and preparing prompts, tools and workflows for production.
An OpenAI essay argues that advanced AI may matter most by speeding up the routine execution work behind breakthroughs, which could set the pace of economic progress.
Customer story: grocery retailer Albertsons Companies describes using ChatGPT Enterprise and the OpenAI API to speed up internal work and make shopping easier for customers.
Customer story: social club The Den says ChatGPT Work shortened tasks such as grant applications from days to hours, saving its team 10–15 hours a week.
OpenAI says it disrupted a coordinated attempt to extract protected model reasoning through distillation and is strengthening its defenses against such attacks.
OpenAI is partnering with America's SBDC to expand hands-on AI training and local support for small businesses, alongside a report on how small teams use AI.
ServiceNow AI describes AutoSynthData, a method for turning an enterprise agent's observed weaknesses into many new, verifiable training tasks that fit the target environment.
Hugging Face introduces the Open TTS Leaderboard to make evaluation of multilingual text-to-speech and voice-cloning models more standardized than today's fragmented, arena-based comparisons.
NVIDIA presents Kumo Tabular, a tabular foundation model that predicts new rows through in-context learning instead of training a separate gradient-boosted model for each task.
Multiverse Computing explains its ProvenanceGuard paper, which checks not only whether an MCP agent's claim is true but whether it is attributed to the right source.
H Company released Holo4, agentic computer-use models in 27B dense and 35B-A3B mixture-of-experts sizes on its H Models API, plus an updated Holotron4 Nano.
Mistral AI is opening a Munich hub for physics AI and industrial AI research, working with partners from German industry.
NVIDIA presents a 64GB DGX Spark as a way for developers to build and run increasingly capable open models and AI agents locally.
NVIDIA says OpenAI's GPT-6 Astra Ultrafast runs on NVIDIA Blackwell GPUs; the model is available in the OpenAI API and to eligible ChatGPT Work and Codex users.
NVIDIA makes its case for AI-factory returns, citing a cost of roughly $60 million per megawatt and arguing operators need a clear ROI before committing capital.
NVIDIA opened applications for its 2027–2028 Graduate Fellowship program, with awards of up to $60,000.
NVIDIA and CoreWeave describe their long co-engineering partnership and an AI-focused cloud covering agentic AI from training to production.
Ai2 released AstaBrief, an 8B open-weights model that writes cited scientific reports; it powers Asta's Fast mode and can be downloaded to run on your own infrastructure.
Ai2 introduced Olmo-core 3, a redesigned and fully open training stack for scaling mixture-of-experts models toward the trillion-parameter range.
MIT researchers present InstructMesh, a tool that produces easy-to-edit designs of everyday objects so both experts and beginners can repair AI-generated 3D models and fabricate them.
MIT's Transit Lab will build an open-source Public Transit Intelligence Hub, funded with $2.1 million from Google.org, to unify transit monitoring, operations and rider communication.
MIT News reports on an AI system that beats top-ranked human Stratego players and is more efficient than other models; the researchers see uses in strategic decision-making.
An MIT study of hiring decisions finds that many firms relying on the same algorithm can, in some situations, benefit job seekers.
MIT professor Sherry Turkle's new book, "Artificial Intimacy," critiques chatbots and the antisocial dynamics she argues they encourage.
Stories with AI-related titles, ranked by points over the past 36 hours.
Automatic selection of new submissions, preferring papers cross-listed in cs.AI, cs.CL and cs.LG. Not a quality ranking.
We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures.
When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it.
Structured tool calls often fail after only a small number of fields violate a schema or an execution contract.
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback.
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused.
Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed.
We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer.
Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment.
In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates.
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps.
Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives.
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training.
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation.
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end.
Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency.
Not included:
Updated: 2026-10-02 16:45 UTC