Our approach to EU text provenance rules
OpenAI describes how it is approaching text watermarking under EU rules: where watermarks apply, how detection works, and why access to detection starts with researchers.
ข่าว AI รายวันจากแหล่งทางการ Hacker News และ arXiv
ชื่อเรื่องและสรุปในหน้านี้แปลด้วยเครื่องจากภาษาอังกฤษ (ออฟไลน์) และอาจมีข้อผิดพลาด โปรดยึดต้นฉบับที่ลิงก์ไว้
การโฆษณาภายใน ChatGPT ยังคงเติบโต: OpenAI จะเพิ่มรูปแบบโฆษณาภาพรวมถึงการวัด การกําหนดค่าและเครื่องมือการสอดคล้องแบรนด์ สิ่งที่สําคัญสําหรับนักตลาดและผู้ใช้ประจําวัน
OpenAI
A concrete look at how one major lab plans to handle EU text-provenance rules, including where watermarks apply and why detection access starts with researchers.
OpenAI
กระดาษที่หายากที่ตรวจสอบหัวข้อของตัวเองเรียกร้อง: ผู้เขียนรายงานว่าการแก้ไขข้อผิดพลาดในการวิเคราะห์ของพวกเขาลดการประหยัดค่าใช้จ่ายที่เรียกว่า 23.9% ไปยัง 4.3% การระมัดระวังที่เป็นประโยชน์สําหรับทุกคนที่ประเมินโมเดล routing ตัวแทน
arXiv
หมายเหตุบรรณาธิการ: สรุปเหล่านี้ร่างขึ้นโดยใช้ AI ช่วยจากข้อความของผู้เผยแพร่แต่ละราย และยังรอการตรวจทานโดยมนุษย์
โพสต์จากบล็อกทางการของบริษัทและห้องวิจัยที่เผยแพร่ใน 7 วันที่ผ่านมา
OpenAI describes how it is approaching text watermarking under EU rules: where watermarks apply, how detection works, and why access to detection starts with researchers.
ต้นฉบับ: Building advertising for the way people use AI
OpenAI เปิดตัวรูปแบบโฆษณาภาพใหม่ใน ChatGPT และเครื่องมือวัดที่ขยาย การพันธมิตรการอนุมัติและการควบคุมการสอดคล้องกับแบรนด์สําหรับผู้โฆษณา
ทั้งหมดของ AI Together Link นํารุ่นเปิดเช่น GLM 5.3 และ Kimi K3 ในทีมเข้ารหัสตัวแทนที่ใช้งานอยู่แล้ว ร่วมกันบอกว่ามันสามารถลดค่าใช้จ่ายรุ่นได้มากกว่า 50%.
เรื่องที่มีชื่อเกี่ยวกับ AI เรียงตามคะแนนในช่วง 36 ชั่วโมงที่ผ่านมา
ต้นฉบับ: OpenAI "rogue" agent activities found on Wikimedia projects
ต้นฉบับ: Homa: The end of TCP for AI clusters [video]
ต้นฉบับ: Spending on AI Is Becoming Almost Impossible for Businesses to Budget
ต้นฉบับ: Accept 'bad things' in return for benefits of AI, says Sam Altman
ต้นฉบับ: Florida woman arrested for allegedly making threats in an AI chat
ต้นฉบับ: People are asking ChatGPT to help them decide how to vote in the midterms
ต้นฉบับ: Building a RAG pipeline for semantic code search
คัดเลือกบทความใหม่โดยอัตโนมัติ (ให้ความสำคัญกับบทความที่อยู่ในทั้ง cs.AI, cs.CL และ cs.LG) ไม่ใช่การจัดอันดับคุณภาพ
การประเมินแบบครบวงจรของรูปแบบการตัดสินใจแบบเปิดน้ําหนักและรูปแบบการตัดสินใจแบบเดียวที่โฮสต์ผ่าน 11 จุดการตัดสินใจของตัวแทนพบว่ารูปแบบที่โฮสต์มีความแม่นยํามากขึ้นใน 9 แห่ง ไม่ได้ชนะโอกาสบนแบบกําหนดเอง Zero-shot และผู้เขียนแก้ไขหลายข้อผิดพลาดในการวิเคราะห์ของตัวเอง
ต้นฉบับ: Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
Auditing 33 LLM-judge configurations against 45,796 worker ratings, the authors find judges can rank responses reasonably well yet estimate acceptance rates anywhere from 3.0% to 97.9%, versus 61.1% for matched workers.
Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy.
ต้นฉบับ: FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training
Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget.
สิ่งที่ทําให้เรื่องราวสั้นกังวล บทความข่าวที่คุ้มค่า หรือการพิสูจน์แม่นยํา elegant?
Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration.
ต้นฉบับ: Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case
This paper studies when LLM investigators should close a case, using 731 audited accident, defect and outage cases; it reports that even a frontier model overstated its evidence in 91% of answers.
ต้นฉบับ: Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs
Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen.
ต้นฉบับ: FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms
ห้องแชททางการเงินหลายฝ่ายเป็นสิ่งสําคัญสําหรับผู้เชี่ยวชาญด้านการขายและเทรด แต่ความซับซ้อนของพวกเขาทําให้การกู้คืนด้วยตนเองของเทรดที่หายไปไม่สามารถทําได้: ทุกคําขอสําหรับใบเสนอราคา (RFQ) เป็นเหตุการณ์ที่ราคาสุดท้ายและผลการค้าปรากฏขึ้นหลายข้อความหลังจาก...
ต้นฉบับ: APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory
APDMem organises a long conversation history into four layers, from thematic summaries down to raw messages, and drills deeper only when a query needs it; on LongMemEval it reports strong results while reading about 8% of the conversations.
ต้นฉบับ: Lexicographic Multi-Objective On-Policy Distillation
การเรียนรู้การเสริมสร้างจากรางวัลที่สามารถตรวจสอบได้ (RLVR) โดยปกติจะเพิ่มความถูกต้องของคําตอบ อย่างไรก็ตามพฤติกรรมแบบจําลองภาษาที่มีประโยชน์ยังต้องการความคิดที่มีคุณภาพสูงและตอบสนองที่เข้มงวด
ต้นฉบับ: Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval
รูปแบบภาษาขนาดใหญ่ (LLMs) ได้แสดงประสิทธิภาพที่แข็งแกร่งในการสร้างรหัส เมื่อความสําเร็จขึ้นอยู่กับทั้งการจดจําความรู้อัลกอริทึมที่เกี่ยวข้องและคํานวณเกี่ยวกับวิธีการประยุกต์ใช้
ต้นฉบับ: How Causality Bridges the Semantic Gap
การวัดดิจิตอล capture how a system behaves แต่มักจะทําให้ความหมายของการเปลี่ยนแปลงของมันไม่ระบุ
LEAP speeds up LLM agents by training a small drafter model on the target model's own action sequences so its proposed actions match more often, guided by a latency framework for speculative action rounds.
ต้นฉบับ: OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
งานหลายภาษาที่เป็นประโยชน์แบบจําลองไม่สามารถประเมินได้โดยการตรวจสอบผลลัพธ์ที่แม่นยํา
ไม่ได้รวม:
อัปเดต: 2026-10-05 20:08 UTC