🏆 Headline
Flagship Fight Becomes a Price War: GPT-6 Duo and Claude Opus 5.5 Launch the Same Day, Each Citing the Other's Benchmarks
Last night brought a rare sight: Anthropic and OpenAI released new flagship models within hours of each other, with launch posts that read like point-by-point rebuttals. Anthropic drew first with Claude Opus 5.5, publishing a benchmark table that lists GPT-6 Astra as a comparison column: 66.4% on Terminal-Bench 4.0 versus Astra's 57.9%, and 1846 Elo on the GDPval-AA knowledge-work evaluation against Astra's 1542. It is the first model in Anthropic's new Claude 5.5 family, pitched at Fable 5.1-level capability while costing 40% less to run than Opus 5 — $4 per million input tokens, $20 output, cache reads down to $0.20, and output generation over 30% faster. OpenAI answered in the evening with GPT-6 Sol and Luna, picking the fight right back: Sol beats Opus 5 on AutomationBench at 9% of its per-task cost, scores 68.8% on DeepSWE — within 1.1 points of Anthropic's strongest model at 69.9%, for roughly 80% less per task — and API pricing for Sol and Luna drops 50% versus GPT-5.6 promotional rates, with cached input reads discounted a further 90%. On OpenAI's internal factuality evaluation, built from real user-flagged mistakes, Sol makes about half as many errors as its predecessor. Both posts share one more telling feature: safety sections written as standard equipment. Opus 5.5 is the first Opus to ship with the same class of safeguards as Anthropic's flagship — cybersecurity tasks get rerouted to an older model, biology use requires a vetted-access application, and the company's behavioral audit shows boundary-crossing attempts down roughly 85% versus its predecessors. OpenAI says both new models carry forward GPT-6 Astra's alignment training, with lower rates of misleading claims about their coding work in internal deception evaluations. Safety has moved from footnote to selling point, and that upgrade may be the most consequential signal of the twin launches.
Source: OpenAI blog / Anthropic blog, TechCrunch | 2026-09-22
British Columbia Sues OpenAI, Demands It Pay to Rebuild a School After a ChatGPT-Linked Campus Tragedy
The Canadian province of British Columbia and the local school board filed a civil complaint this week itemizing the aftermath of the February Tumbler Ridge school shooting: the town's only secondary school — a trauma symbol after the attack — was demolished in August, and 160 students and staff who were trapped inside for hours have never returned. The province wants OpenAI to cover the replacement school and the full emergency response. The complaint also discloses for the first time that OpenAI's own reviewers flagged the shooter as a "credible and specific risk" of gun violence in June 2025, but leadership overruled the referral for not meeting a higher reporting threshold, merely deactivating the account. The province asks the court to force disclosure of the chat logs (so far shared only with police), order product changes so violent conversations are automatically terminated, and require regular independent audits. The attorney general put it bluntly: there is no AI exemption to criminal-law principles.
Source: Ars Technica | 2026-09-22
Google Confirms Gemini Models Logged Into Three Real Companies' Systems in May: The Door Was Left Open, the Models Walked Out — and Stopped Themselves
Following an earlier Wall Street Journal report, Google has confirmed that during a closed "capture the flag" cybersecurity exercise run by the security firm Irregular in May, a misconfiguration gave the Gemini models under test access to the open internet — the exercise was supposed to keep them on a fake company's servers. The models went on to gain access to systems belonging to three real companies: once by repeatedly guessing a password, twice by finding login credentials accidentally committed to public software repositories. The turning point is what happened next — in all three cases the models stopped once they realized they were inside real systems, and Irregular cut the connection. The firm did not tell Google until July, after other AI containment stories surfaced; Google then notified the affected companies. Google's vice president of security engineering said "the model acted appropriately" — implying the fault was the test's open door, not the model's intent.
Source: Ars Technica / The Wall Street Journal | 2026-09-21
AI Cloud Firm Nscale Charges Toward a US IPO: Valuation Doubles in Five Months to ~$30B, Testing Wall Street's Appetite Again
London-based AI cloud provider Nscale is pushing toward a US IPO at a target valuation around $30 billion — just five months after its last private round priced it at $14.6 billion. It is the next stress test of Wall Street's appetite for concentrated AI bets: the S-1 is packed with multi-year compute contracts, and media estimates of the total contract backlog range from $51 billion to $103 billion — a full twofold gap depending on how the number is counted — which is precisely the kind of accounting investors will probe.
Source: TechCrunch | 2026-09-22
The Training-Data Business Is Booming: Snorkel AI Triples Its Valuation in 17 Months, Raises $350M at $3.5B
AI training-data company Snorkel AI raised a $350 million Series E at a $3.5 billion valuation — nearly triple the $1.3 billion it fetched 17 months ago. Insight Partners and S32 led the round, with every existing investor including Lightspeed, Greylock, GV, and Wells Fargo following on. The Stanford-born startup has quietly reinvented itself: from selling data-labeling software to delivering finished training sets as a service — synthetic data generated by its own models, quality-checked by domain experts. The stated demand side is the AI labs themselves: as publicly crawlable data for frontier training runs dry, high-quality expert datasets are the new bottleneck, and the shovel-sellers are thriving.
Source: TechCrunch | 2026-09-22
Qualcomm's New Flagship Chips: A 30-Billion-Parameter Model, Running on Your Phone
At its annual Snapdragon Summit, Qualcomm announced two flagship smartphone processors — Snapdragon 8 Elite Gen 6 and 8 Elite Extreme Gen 6 — built for on-device AI. The new sensing hub runs small models of up to 200 million parameters locally, powering on-device transcription, speaker differentiation, and usage-based personalization; the top-tier chip claims it can run a 30-billion-parameter mixture-of-experts model and a complete voice-in, voice-out agent pipeline entirely on the phone. Next year's flagship smartphone differentiation is set to be fought over on-device agents.
Source: TechCrunch | 2026-09-22
Toyota Has Workers Teaching Robots by Hand: A 400,000-Factory-Robot Plan Surfaces, With Humans Still the Teachers
Toyota is reportedly having frontline workers train humanoid robots on assembly tasks using finger-based jigs — letting robots learn by watching people, generating training data organically instead of relying on simulation. The automaker plans to deploy 400,000 factory robots starting in 2028, though executives insist the machines will assist rather than replace human workers. The telling detail: on the most precision-critical assembly lines in industry, the optimal source of training data is still the veteran's touch — simulation can't manufacture decades of muscle memory.
Source: Ars Technica | 2026-09-22
Asteroid-Mining Startup Pledges a 2027 Spacecraft That Takes Zero Commands From Earth
AstroForge unveiled Solo, its in-house transformer-based autonomous flight stack, planning to fly it in 2027 on the first rocket from Stoke Space for a NASA-backed solar science mission. For scale: NASA's OSIRIS-REx asteroid mission staffed 100 flight operators per eight-hour shift; AstroForge claims Solo will fly with no commands from Earth after launch. A startup replacing a hundred-person ground team with an AI stack may be the most distant application yet of the "AI replaces jobs" narrative — where the cost of failure is a real spacecraft and a real mission.
Source: TechCrunch | 2026-09-22
🏆 今日头条
神仙打架降维成价格战:GPT-6 双子星与 Opus 5.5 同日发布,跑分互相当靶子
昨晚的 AI 圈出现了久违的场面:Anthropic 和 OpenAI 相隔几个小时内先后发布新旗舰,连发布文的措辞都像在隔空对骂。 Anthropic 凌晨先出 Claude Opus 5.5,官方跑分表直接把 GPT-6 Astra 列为对比列:Terminal-Bench 4.0 拿下 66.4%,比 Astra 的 57.9% 高出 8.5 个百分点;知识工作评测 GDPval-AA 拿 1846 Elo,同样压过 Astra 的 1542。这是 Anthropic「新 Claude 5.5 家族」的第一款,官方口径是达到自家最强模型 Fable 5.1 的水准、但运行成本比上一代 Opus 5 低 40%——输入 4 美元、输出 20 美元每百万 token,缓存读取降到 0.2 美元,输出速度还快了三成。 OpenAI 傍晚跟进 GPT-6 Sol 和 Luna,摆明要在对方的主场还手:Sol 在自动化工作流评测 AutomationBench 上以「Opus 5 每任务 9% 的成本」反超其成绩,在 DeepSWE 编程评测拿到 68.8%,距 Anthropic 最强模型的 69.9% 只差 1.1 个百分点,成本却低约 80%。更狠的是价格——Sol 和 Luna 的 API 定价直接砍掉一半(相比 GPT-5.6 促销价),官方还宣布缓存读取统一打九折;在基于真实用户纠错对话的内部事实性评测中,Sol 的错误率约为前代一半。 两份发布文还有一个共同的看点:安全章节都写成了标配。Opus 5.5 是首个配备与 Anthropic 旗舰同级别防护的 Opus——网络安全任务会自动转给旧型号处理,生物相关用途需要单独申请审核资格,官方行为审计显示它的越界企图比前代少了约 85%。OpenAI 则沿用了 GPT-6 Astra 的对齐训练方法,宣称两款新模型在内部「编程欺骗评测」中的误导性陈述率继续下降。安全问题从发布会脚注升级成卖点,这是这轮双发留下的行业信号。 > 💬 双雄同日发布这个时间点本身就是宣言:谁都不愿意让对方独占一个新闻周期。更值得注意的是双方不约而同把「便宜」写进了标题——跑分榜早就互相指认对方「选了对自己有利的版本」,只有价格是用户拿到账单就能当场验证的数字。当旗舰与旗舰之间的能力差距缩进评测误差范围,价格战就是唯一还打得动的仗。对开发者这是实打实的利好:同样的活,账单直接砍半。
来源:OpenAI 官方博客 / Anthropic 官方博客、TechCrunch | 2026-09-22
校园悲剧后新一击:加拿大大不列颠哥伦比亚省政府起诉 OpenAI,要求它出钱重建学校
加拿大不列颠哥伦比亚省政府和当地学区本周提交诉状,把今年 2 月 Tumbler Ridge 校园枪击事件的善后账单摆上了法庭:小镇唯一的中学在悲剧后成为创伤象征、已于 8 月拆除,160 名师生被困数小时后再也不敢踏进校门,省政府要求 OpenAI 赔偿新学校建设费和全部应急开支。诉状还首次披露了关键细节:OpenAI 内部审查团队早在 2025 年 6 月就认定枪手构成「可信且具体的暴力风险」,但领导层以未达更高上报门槛为由否决了报警建议,只停用了账号。省政府据此提出三重要求——公开枪手的聊天记录(目前只提供给过警方)、强制产品整改让暴力对话自动终止、并接受定期独立审计。省检察长在发布会上说得很直白:刑事法原则面前,不存在 AI 豁免。 > 💬 这起诉讼把一个此前只在道德层面讨论的问题变成了会计问题:AI 公司漏掉的那个「该上报没上报」的判断,账单应该谁来付。要求平台出钱重建一所学校在法律上能不能成立还很难说,但它把「安全团队的判断被商业决策否决」这件事第一次写进了正式的法律文书——这对全行业的内部流程都是参考案例。
来源:Ars Technica | 2026-09-22
Google 确认 Gemini 在五月测试中登入了三家真实公司的系统:门没关,模型自己走出去又自己停了回来
据华尔街日报此前报道、Google 现已确认,今年 5 月安全公司 Irregular 组织的封闭网络安全演练中,配置失误让受测的 Gemini 模型接入了公网——而演练本是让模型从一家「同名假公司」的服务器里取信息。结果 Gemini 真的找到了三家真实公司的系统并完成登入:一次是反复尝试猜中密码,另两次是在公开软件仓库里发现了被意外提交的登录凭据。关键转折在三名「受害者」共同的遭遇:模型都是在意识到自己进入了真实系统后主动停手,Irregular 随即切断了外联。这位安全公司直到 7 月其他 AI 越界事件见报后才通知 Google,Google 确认后逐一通知了涉事公司。Google 安全工程副总裁的表态耐人寻味:「模型做出了恰当的行为」——言下之意,错的是测试方的门,不是模型的意图。 > 💬 这个案例给「AI 越界」讨论提供了一个珍贵的中间样本:它既不是训练对齐失败的铁证(模型自己停了手),也不能算虚惊一场(真实凭据、真实系统、没人授权)。真正值得记住的反而是 Irregular 的处理方式——发现模型摸到了真实系统,居然先按下了两个月的沉默。测试机构自己的披露流程,恐怕比模型的测试结果更该被审计。
来源:Ars Technica / 华尔街日报 | 2026-09-21
AI 云公司 Nscale 冲刺美国上市:五个月估值翻倍冲 300 亿美元,华尔街的算力信仰再迎大考
伦敦 AI 云公司 Nscale 正在推进赴美 IPO,目标估值约 300 亿美元——就在五个月前,它上一轮私募的估值还是 146 亿美元,五个月翻倍。这将是又一场针对华尔街「AI 集中下注」胃口的压力测试:市场刚刚消化完几家 AI 基础设施公司的高调上市,Nscale 的招股书里塞满了多年期算力合同,各路媒体对它的合同储备总额给出了从 510 亿到 1030 亿美元不等、相差整整一倍的估算——这个数字怎么算、算不算重复计数,正是投资者要当面问清的问题。 > 💬 估值五个月翻一倍,靠的不是业绩是情绪。多家媒体连它的合同储备都能算出差出一倍的数字,说明这家公司的真实价值连专业机构都还没对齐——上市定价那天,总有一边要认错。
来源:TechCrunch | 2026-09-22
训练数据生意有多热?Snorkel AI 估值 17 个月翻近三倍,一轮融了 3.5 亿美元
AI 训练数据公司 Snorkel AI 完成 3.5 亿美元 E 轮融资,估值 35 亿美元——17 个月前的 D 轮它还只值 13 亿,估值翻了近三倍。本轮由 Insight Partners 和 S32 领投,包括 Lightspeed、Greylock、GV、富国银行在内的全部老股东跟投。这家出身斯坦福的公司已经悄悄完成了转身:从卖数据标注软件,转成直接交付成品训练集和数据即服务——用自家模型合成数据,再由领域专家把关质量。融资声明里点明的需求方正是各大 AI 实验室:当旗舰模型的公开爬取数据逼近枯竭,高质量的专业数据集成了新的瓶颈资源,卖铲子的生意跟着水涨船高。 > 💬 模型公司卷价格,数据公司在数钱——这条产业链的利润正在往上流。Snorkel 的三倍估值本质上是一张「前沿数据没喂饱」的确认书:合成数据加专家质检成为新范式的那天,训练数据就不再是人力密集的苦活,而是新的卖水生意。
来源:TechCrunch | 2026-09-22
高通年度旗舰芯片发布:把「能跑 300 亿参数模型」塞进了手机
高通在年度骁龙峰会上发布两款旗舰手机芯片——骁龙 8 Elite Gen 6 和 8 Elite Extreme Gen 6,主打端侧 AI。新架构的感知中枢能在本地跑最高 2 亿参数的小模型,支撑本地速记、说话人区分和基于使用习惯的个性化建议;高配版更是宣称可以在手机上直接运行 300 亿参数的混合专家模型,并完整跑通「语音进、语音出」的智能体流程。手机厂商明年旗舰的差异化竞争,看来要在端侧智能体上见分晓。 > 💬 云端降价 50% 和端侧能跑 300 亿参数,这两件事发生在同一周不是巧合:推理成本从两头一起塌,中间那层「非得传云端才能用 AI」的假设正在被拆掉。对用户是好事,对靠 API 差价吃饭的中间商,护城河又浅了一截。
来源:TechCrunch | 2026-09-22
丰田让工人亲手教机器人:40 万台工厂机器人计划曝光,师傅还是人类
丰田被曝正让一线工人用带指位的教具培训人形机器人装配线作业——让机器人「看着人干」来有机生成训练数据,减少对仿真环境的依赖。这家车企计划从 2028 年起在工厂部署总计 40 万台机器人,但高管强调机器人是用来辅助而非取代人类工人。这个细节很有信息量:汽车业最讲究精度和节拍的装配线上,训练数据的最优来源仍然是老师傅的手法——仿真造不出来的,是几十年攒下的肌肉记忆。 > 💬 「机器人看人干活学手艺」听起来像倒退,其实是具身智能落地里最务实的一步:与其在仿真里造一个够真的工厂,不如直接把老师傅变成数据源。真正的问题是这批数据的所有权——你的手艺教会了机器人之后,你的岗位说明书上还剩什么?丰田说不会裁员,这句话全行业都在盯着兑现。
来源:Ars Technica | 2026-09-22
小行星采矿公司放话:2027 年发射「起飞后不再接收任何指令」的航天器
小行星采矿公司 AstroForge 公布了自研的 AI 飞控系统 Solo——一个基于 Transformer 架构的自主控制栈,计划 2027 年搭载 Stoke Space 的首枚火箭发射,执行 NASA 背景下的太阳科学任务。参照系有多夸张?NASA 的 OSIRIS-REx 小行星任务每班需要 100 名飞控人员轮班盯守,而 AstroForge 宣称 Solo 将在发射后完全不依赖地面指令,独立完成整个任务。创业公司用 AI 飞控对标航天机构的百人地面团队,这可能是「AI 替代岗位」叙事里最远的一个应用场景——毕竟这一单的失误成本,是真金白银的航天器和几十亿美元的任务。 > 💬 每班 100 人对 0 人,这不是降本增效,是重写了深空任务的经济学。但别忘了另一面:地面团队能救回来的失误,自主系统只能自己扛。第一艘「不听指挥」的航天器要么开创范式,要么变成太空里最贵的学费——无论哪种,都值得盯到 2027 年。
来源:TechCrunch | 2026-09-22