📰 AI Frontier Daily

AI Frontier Daily

Keywords: 0% probability, open-weights crown, AI going independent
关键词:0% 概率、开源王座、AI 自立门户
🏆 Headline

Jensen Huang Rates Extinction Risk at "0%": Refusing Panic — and Dismantling the Case for New Rules

Nvidia CEO Jensen Huang took the rhetoric to its limit in a CBS News interview: "2030 is not going to be the end of the world. There is 0% chance that's going to be the end of the world. Scaring people is unnecessary. It is irresponsible." According to The Guardian on September 21, he directly criticized former Anthropic researcher Jacob Coxon's claim that AI could become "superhuman" and kill off humanity within the decade — Coxon and two other Anthropic researchers had put the probability of AI destroying humanity above 10%, sparking an international debate over AI's dangers and the lack of oversight. But Huang's real novelty came in the second half: he argued that AI companies' recent calls for national and international safety regulation are actually aimed at being "relieved of existing laws" — "Go and read between the lines. They're actually not asking for more laws. They're asking to be relieved of the laws we do have." He cited cyberUnauthorized-entry laws and damage liabilities: "If your product does harm, if your product doesn't perform as you promised, there are plenty of laws. Apply those first." The timing was pointed: the four labs that called for a slowdown were just sued by consumers over an alleged coordination pact (covered in our September 20 edition), and this 0%-versus-10% wager puts the industry's deepest safety schism squarely on the table.

Source: The Guardian / CBS News | 2026-09-21

OpenAI Urges the US to Lead Global AI Safety Standards, With "Self-Improvement" in the Crosshairs

OpenAI published a policy blog on Monday urging the US government to lead a multi-country effort on global technical standards for frontier AI. Per Politico and Bloomberg, the most striking item on the standards list is recursive self-improvement — models improving models — alongside model evaluation methods and incident reporting. The company once criticized for building AGI behind closed doors is now handing regulators a technical roadmap. Timing matters: the UN Security Council is debating AI risks this week, in the same week Huang declared existing laws sufficient. When a lab starts drafting the regulators' checklist for them, read both messages at once: it may be genuinely afraid, and it may want to define "how to be regulated" before being regulated badly. Both are probably true — that's what makes it interesting.

Source: Politico / Bloomberg | 2026-09-21

Xiaomi's MiMo-V2.6 Takes the Open-Weights Crown: 7 Points From the Closed Frontier, Full RL Pipeline Published

Xiaomi released MiMo-V2.6 in one stroke: Pro and Flash, both omnimodal. Pro scored 46 on the Artificial Analysis Intelligence Index — per Artificial Analysis and OfficeChai, that puts it atop the open-weights leaderboard: the crown had been shared at 44 by Z.AI's GLM-5.3 and Moonshot's Kimi K3, and Xiaomi's 20-point generational jump (its predecessor MiMo-V2.5-Pro sat at just 26) takes the title, ranking sixth overall — 7 points behind co-leaders Claude Fable 5.1 and GPT-6 Astra at 53. Architecturally, Pro is a mixture-of-experts model with 1.02 trillion total parameters and 42 billion active. The more notable part is how it was open-sourced: alongside the weights, Xiaomi published the technical report, its RL environments, and the training code — handing over "how it was trained" together with "what was trained." By Artificial Analysis' accounting it costs $0.13 per Intelligence Index task, landing on the intelligence-versus-cost Pareto frontier.

Source: Artificial Analysis / OfficeChai | 2026-09-21

OpenAI Convenes a Nine-Member Math Advisory Group: Terence Tao's Blog Announces the Task — Scheduling the Discharge of 100+ New Theorems

OpenAI announced it is working with an independent Advisory Group on Mathematics and Artificial Intelligence. Per its official blog: a new internal model that began training on August 28 has, after resolving the Navier–Stokes Millennium Prize problem, now resolved more than 100 long-standing open problems across most areas of mathematics — progress that "has surprised the mathematicians within OpenAI." The group is hosted at the Institute for Advanced Study in Princeton, and the nine initial members read like a living all-star team: Fields medalists Timothy Gowers and Martin Hairer, plus Edward Witten, Melanie Matchett Wood, and others. Terence Tao's blog carried the group's guest announcement on Monday: it operates independently, members are unpaid by OpenAI, and it can make its advice public — "the very specific challenge" being "advising OpenAI on how to coordinate the release" of the results produced by its internal model. The comment section erupted: one PhD student calculated that four years of doctoral work might not survive an agent swarm's day; others warned the group is crisis management for OpenAI's IPO. The group explicitly has no mandate over the pace of OpenAI's mathematics research.

Source: OpenAI Blog / Terence Tao's Blog | 2026-09-21

Googlebook Preorders Open: From $899, Betting You'll Switch Laptops for Gemini

Per Ars Technica and TechCrunch, Google confirmed Monday that its Googlebook laptops go on sale October 4 starting at $899, with preorders for five models live. This is not a Chromebook refresh: the devices run Android OS with a desktop Chrome browser, and the real pitch is Gemini in laptop form — an AI-powered cursor, vibe-coded widgets, and a dictation feature called Rambler that cleans messy spoken brain-dumps into readable text. The wager is legible: Google is betting "AI-exclusive features" can become a reason to switch hardware, the way it once bet "the browser is enough." The last bet with this shape was called the netbook. This time the desktop has already been colonized by agents — whether the laptop is the next entry point gets its answer on October 4.

Source: Ars Technica / TechCrunch | 2026-09-21

Amazon Shuts the Door on Meta's Muse: Agentic Commerce Meets Platform Sovereignty

Starting Sunday night, Muse users trying to check out on Amazon began seeing a new error: "Continued access by an unauthorized AI agent violates Amazon's Conditions of Use." As spotted by GeekWire and covered by TechCrunch, Muse is now explicitly blocked from amazon.com. TechCrunch adds a cooler-headed layer: this isn't just giant-on-giant sparring — Amazon has its own foundation models and one of the most popular inference platforms on the internet, and with no legal obligation to open the door, why would it? The operational logic is even simpler: if Muse makes a bad order, Amazon cleans it up, appeasing both the angry customer and the angry vendor. Muse's hallucination rate is among the lower ones, but still far from zero. Lesson one of agentic commerce, now written: platforms owe agents no entrance.

Source: TechCrunch / GeekWire | 2026-09-21

Muse's Numbers Keep Climbing: Downloads and Daily Actives Both Above ChatGPT's Early Trajectory

The good days continue. Per TechCrunch citing fresh Apptopia estimates: in its first 12 days, Muse hit 1.8 million downloads on iOS in the US and Canada, versus 1.3 million for ChatGPT over the same window; 2.8 million installs globally. The daily-active gap is starker — 642,000 US mobile DAU for Muse against 231,000 for ChatGPT at the same point. Even narrowing the comparison to iOS alone (Muse is on both stores; ChatGPT launched iOS-only), Muse leads with 359,000 DAU. The app climbed from No. 2 overall on the US App Store at launch to No. 1. One background figure deserves attention: Apptopia found over 95% of Muse users are also Facebook users, and 63% are on Instagram — the "cold start" was a hot start fueled by Meta's own traffic pools. Meta has yet to publish any official Muse figures.

Source: TechCrunch / Apptopia | 2026-09-21

Grok 4.7 Launches: Half the Price of the Frontier, Benchmarks Expose the Other Side of "Value"

SpaceXAI released Grok 4.7 overnight, pitching best-in-class price-performance: $2 per million input tokens and $6 per million output tokens — half of the comparable models. Official numbers show gains concentrated in long tasks and specialized domains: CursorBench 4.0 up from 40.4% to 46.3%, EEBench from 53.0% to 64.0%, and 19.6% on the Harvey Legal Agent Benchmark — nearly triple its peer. But The Decoder poured cold water with third-party data: on Terminal-Bench 4.0, independent testing puts Grok 4.7 at just 26%, versus roughly 60% for GPT-6 Astra and 55% for Claude Fable 5.1 — a 12-point gap from xAI's own 38% figure. The price war is real. So is the war over who interprets it.

Source: xAI Blog / The Decoder | 2026-09-21

🔧 Recommended Tools

ToolTypeHighlight
[kev](https://github.com/jaredpalmer/kev)Decision modelsA tiny Jev-like family of decision models built on top of Qwen3.5 — train and run it yourself; hit 384 points on HN overnight (GitHub 2,377★, within five days)
[typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use)Desktop automationComputer use at about $0.0002 a step: OCR the screen, classify the next action with TypeSafe, click — native on macOS (GitHub 747★, within six days)
[mini-AGI](https://github.com/volotat/mini-AGI)Continual learningA continual-learning model trained from scratch on an 8GB-VRAM laptop: batch size 1, streaming data — the minimalist sample of continual learning on personal hardware (GitHub 298★, within three days)
🏆 今日头条

黄仁勋给灭世论打出「0%」:不仅拒绝恐慌,还要拆掉新规的桌子

Nvidia CEO 黄仁勋在接受 CBS News 采访时把话说到了极致:「2030 年不会是世界末日,这种事发生的概率是 0%。吓唬人没有必要,是不负责任的。」据 The Guardian 9 月 21 日报道,他直接点名批评前 Anthropic 研究员 Jacob Coxon「AI 可能在十年内变得超人类并消灭人类」的说法——后者与另外两名 Anthropic 研究员此前声称 AI 毁灭人类的概率超过 10%,这番言论在国际上掀起了一场关于 AI 危险性与监管缺位的大辩论。但黄仁勋真正的新意在后半段:他认为 AI 公司近来呼吁的国内国际安全监管,目的其实是「豁免于现有法律」——「你们去读读字里行间,他们要的不是新法律,而是免除我们已有的法律。」他点名网络安全领域的未授权访问法和损害赔偿责任:「如果你的产品造成了伤害,如果产品没有兑现承诺,现有法律管得够多了。先把这些法律用起来。」这番表态来得恰是时候:四家实验室刚被消费者以「合谋降速」告上法庭(本报 9 月 20 日报道过这桩诉讼),而这场 0% 与 10% 的对赌,等于把行业最大的安全分歧摆上了明面。 > 💬 这场争论最讽刺的地方在于:一边说「毁灭概率 10%」的公司在喊监管,一边说「概率 0%」的公司反而不信监管——黄仁勋的逻辑其实自洽:如果真是 0%,现有法律当然够用;但如果大家都真信自己的数字,这行业就不会有这么多话事人了。两个数字都对不了账,但两个数字都在定价。

来源:The Guardian / CBS News | 2026-09-21

OpenAI 呼吁美国牵头制定全球 AI 安全标准,靶心对准「自我改进」

OpenAI 周一发布政策博客,呼吁美国政府牵头与多国共同制定前沿 AI 的全球技术标准,据 Politico 与 Bloomberg 报道,标准清单里最扎眼的一条是「递归自我改进」(recursive self-improvement)——也就是模型改进模型本身的能力——此外还包括模型评估方法与事件报告机制。这家曾被批评「闭门造 AGI」的公司,如今主动给各国监管者递上了技术路线图。值得注意的是时机:UN 安理会本周正在讨论 AI 风险,而就在同一周,黄仁勋说现有法律够用了。实验室自己出题、自己划重点,监管者接不接这份「作业」,是接下来一年的看点。 > 💬 当一家公司开始替监管者起草监管对象清单,你要同时读出两层意思:它可能真怕,它也可能想先定义「怎么管」来避免「被乱管」。两句话都是真的,这才有意思。

来源:Politico / Bloomberg | 2026-09-21

小米 MiMo-V2.6 开源登顶:离闭源前沿只差 7 分,RL 训练管线全套公开

小米一口气放出 MiMo-V2.6 的 Pro 与 Flash 两个全模态版本,其中 Pro 在 Artificial Analysis 智能指数上拿到 46 分——据 Artificial Analysis 与 OfficeChai 报道,这使它直接登顶开源权重模型榜:此前开源王座由智谱 GLM-5.3 与月之暗面 Kimi K3 以 44 分共享,如今小米以 20 分的代际跨越(上一代 MiMo-V2.5-Pro 仅 26 分)夺下头名,在全模型榜上排第六,距离榜首 Claude Fable 5.1 与 GPT-6 Astra 的 53 分只差 7 分。架构上 Pro 是 1.02 万亿总参数、420 亿激活的 MoE 模型。更难得的是开源的开法:权重之外,技术报告、强化学习训练环境和训练代码一并公开——把「怎么练出来的」和「练出来的东西」一起交了出去。按 Artificial Analysis 的口径,它每任务成本仅 0.13 美元,落在智能/成本帕累托前沿上。 > 💬 开源阵营换庄的速度已经快过闭源的版本号:三个月前还是「Kimi 追平 GPT」,这周就轮到小米。中国实验室之间的王座之战,成了开源世界效率最高的引擎。

来源:Artificial Analysis / OfficeChai | 2026-09-21

OpenAI 拉起九人数学顾问团:陶哲轩博客发文,任务是给 100 多个新定理「安排出院」

OpenAI 宣布与一个独立的「数学与人工智能顾问组」合作,据其官方博客披露:8 月 28 日开始训练的新内部模型,在解决纳维-斯托克斯千禧年大奖难题之后,已解决超过 100 个横跨数学大多数领域的长期悬而未决的问题——进展速度「连 OpenAI 内部的数学家都感到惊讶」。为此成立的顾问组设于普林斯顿高等研究院,九位初始成员名单堪称在世数学家的顶配:菲尔兹奖得主 Timothy Gowers、Martin Hairer,加上 Edward Witten、Melanie Matchett Wood 等。陶哲轩的博客周一同步发了顾问组的客座公告:团队独立运作、成员不从 OpenAI 拿钱、可以公开建议,「当前的具体挑战,是就如何协调发布这批由内部模型产出的重要数学结果,向 OpenAI 提供建议」。博客评论区已经吵翻了:有博士生算出「我四年寒窗可能抵不过 agent 集群一天」,也有人警告这是在替 OpenAI 的 IPO 做危机公关。顾问组同时明确:无权过问 OpenAI 数学研究的推进节奏。 > 💬 请谁当顾问不重要,顾问的权限边界才重要:能管「怎么公布」,不能管「跑多快」。人类数学界拿到的这份合同,签的是发布流程的编辑权,不是科学议程的方向盘——但比起一年前连门都敲不开,这已经是能谈到的最好价格了。

来源:OpenAI 官方博客 / Terence Tao 博客 | 2026-09-21

Googlebook 今晚开启预购:899 美元起,赌你愿意为 Gemini 换一台笔记本

据 Ars Technica 与 TechCrunch 报道,Google 周一确认 Googlebook 笔记本 10 月 4 日正式开卖、899 美元起,五款机型预购通道已经打开。这不是 Chromebook 的改款:设备跑 Android OS 配桌面版 Chrome 浏览器,真正的主打货是 Gemini 的笔记本形态功能——AI 光标、vibe-coded 小组件,以及一个叫 Rambler 的语音输入功能,专门把语无伦次的口述整理成可读文本。这一局的赌注意味深长:Google 在赌「AI 独占功能」能成为换机理由,就像当年赌「浏览器够用了」一样。上一次这个赌局叫上网本,这一次桌面端已经被 agent 吃掉半壁,笔记本是不是下一个入口,10 月 4 日见分晓。 > 💬 硬件厂商终于承认了:大家换电脑不是因为旧电脑慢,是因为旧电脑装不下新模型。899 美元买的不是配置,是一张 AI 桌面端的船票——问题是,手机端的 AI 已经把船开走了。

来源:Ars Technica / TechCrunch | 2026-09-21

亚马逊对 Meta Muse 关门:AI 代购第一次撞上平台主权

上周日晚间开始,Meta 的 AI 助手 Muse 用户在亚马逊结账时收到一条新错误信息:「未授权 AI agent 持续访问违反亚马逊使用条件。」据 GeekWire 首先发现、TechCrunch 跟进报道,Muse 被明确挡在了 amazon.com 门外。TechCrunch 的分析给了一层更冷静的解读:这不只是巨头互掐——亚马逊有自己的基础模型和全美最大的推理平台之一,在没有任何法律义务给 Muse 开门的情况下,为什么要开?更现实的原因是运营责任:Muse 下错一单,收拾残局的是亚马逊——要同时安抚愤怒的买家和卖家。Muse 的幻觉率在模型里已经算低的,但离零还很远。Agent 商务的第一课就此写好:平台不欠 agent 一个入口。 > 💬 这条错误信息是 agent 时代的第一份「平台主权宣言」:你的 agent 再聪明,到了我的地盘就是未授权访客。代购这门生意,恐怕最后还得靠分成协议,而不是浏览器自动化。

来源:TechCrunch / GeekWire | 2026-09-21

Muse 数据继续开挂:下载量与日活全面超过当年 ChatGPT

Muse 的好日子还在继续。据 TechCrunch 援引 Apptopia 最新估计:上线 12 天,Muse 在美国与加拿大的 iOS 下载量达 180 万次,同期 ChatGPT 是 130 万;全球总安装量 280 万。日活差距更夸张——美国区 Muse 日活 64.2 万,当年同期 ChatGPT 只有 23.1 万。即便把对比口径收窄到 iOS 单平台(Muse 双端上线、ChatGPT 当年只有 iOS),Muse 仍以 35.9 万日活领先。应用已从上线时的美国 App Store 总榜第二爬到第一。一个背景数字值得玩味:Apptopia 发现超过 95% 的 Muse 用户同时也是 Facebook 用户,63% 是 Instagram 用户——所谓冷启动,其实是 Meta 用自家流量池完成的热启动。Meta 官方至今未公布任何 Muse 运营数据。 > 💬 「超过 ChatGPT 早期」这个对比有点甜味剂:当年 ChatGPT 是零投放裸奔,Muse 背后站着全球最大的社交分发机器。但反过来说,能把分发优势复用成 AI 入口的,目前全世界也只有这一家——数据是真的,滤镜也是真的。

来源:TechCrunch / Apptopia | 2026-09-21

Grok 4.7 发布:价格砍半对标前沿,跑分却暴露了「性价比」的另一面

SpaceXAI 凌晨发布 Grok 4.7,主打「同类最强性价比」:输入 2 美元/百万 token、输出 6 美元,正好是对位模型的一半。官方数据显示进步集中在长任务与专业领域:CursorBench 4.0 从 40.4% 提到 46.3%,EEBench 从 53.0% 跳到 64.0%,Harvey 法律代理基准 19.6% 是对位模型的近三倍。但 The Decoder 引用第三方数据泼了盆冷水:在 Terminal-Bench 4.0 上,第三方测得 Grok 4.7 仅 26%,而 GPT-6 Astra 约 60%、Claude Fable 5.1 约 55%——xAI 自己报的 38% 与第三方口径差了 12 个点。价格战是真的,价格战的解读权之战也是真的。 > 💬 每次发布会都像两个平行宇宙:官方表格里「全面对标前沿」,第三方跑分里「仍有代差」。这次连 12 个点的口径差都摆在明面上了——买模型前先想清楚你要信谁的表格。

来源:xAI 官方博客 / The Decoder | 2026-09-21

🔧 工具推荐

工具类型亮点
[kev](https://github.com/jaredpalmer/kev)决策模型在 Qwen3.5 之上训练的微型 Jev 风格决策模型家族,可自己训练、本地运行,HN 一夜冲上 384 分(GitHub 2377★,五天内)
[typesafe-computer-use](https://github.com/awlevin/typesafe-computer-use)桌面自动化每步约 0.0002 美元的电脑操作 agent:OCR 读屏、TypeSafe 分类下一步动作、直接点击,macOS 原生(GitHub 747★,六天内)
[mini-AGI](https://github.com/volotat/mini-AGI)持续学习在 8GB 显存的笔记本上从零训练的持续学习模型:批大小 1、数据流式喂入,个人硬件跑持续学习的极简样本(GitHub 298★,三天内)