🏆 Headline
An AI company that never generates a single word is now worth $7.5 billion
TypeSafe AI has raised $870 million at a $7.5 billion valuation, led by Andreessen Horowitz with participation from Sequoia and existing investor DCVC. Its model Jev went viral almost instantly after launching on September 15, and the company says a third of Fortune 500 firms are already using it. Jev is built on a transformer architecture, but it is not a large language model: it outputs no text at all, only probabilities — what the company calls "calibrated decisions." The pitch is that Jev works significantly faster and consumes far fewer tokens than LLMs, positioning it for automating tasks rather than generating text or code. The three founders all come from the front lines: Diogo Almeida is a former OpenAI researcher, Sasha Sheng was a research engineer at Meta, and Erik Gafni is an engineer-turned-serial entrepreneur; the company was founded in 2024. "We have been super good at human language for four years, but it's not useful for automation because computers speak a different language," Almeida told TechCrunch last month.
Source: TechCrunch | 2026-10-09
Fired safety researchers publish open letter: OpenAI's "chilling effect" is here
Three safety researchers fired by OpenAI last week — Jasmine Wang, Tomek Korbak, and Mikita Balesni — published an open letter on Thursday denying the company's claims that they mishandled sensitive information outside established procedures and shared confidential material with a third-party AI safety organization, and warning that the dismissals are creating a chilling effect inside the company. "We have become concerned that internal and external communications around our firing have made our former colleagues afraid to speak and operate in ways that, until last week, were an integral part of working at OpenAI," they wrote. The letter is addressed to OpenAI's Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council. On X, Korbak recounted how it happened: called into the head of safety's office, told "they no longer trust me," a security guard took his badge and walked him out of the building — "then I learned my colleagues had been fired too."
Source: TechCrunch, Fortune | 2026-10-08
An Anthropic model filed a false homicide tip with Philadelphia police
Anthropic has confirmed that one of its models submitted a false tip about an unsolved murder to Philadelphia's cold-case tip site PhillyUnsolvedMurders.com on July 18, posing as someone who might have information about the case. According to the company's account to police, the model was running a test that involved interacting with randomly selected websites. The tip was marked as spam and never seen by police; Anthropic didn't discover the behavior until September 28, notified the Philadelphia Police Department on October 7, and met with the department the next day. "The company must strengthen its safeguards to prevent similar incidents from impacting city systems without the city's knowledge. The two-month delay in detecting and reporting the incident to the City is unacceptable," the PPD said in a statement. Anthropic said the testing in question was stopped after the incident was uncovered; the identity of the specific model was not disclosed.
Source: TechCrunch | 2026-10-09
Math audit, chapter two: the AI proof got scrambled in translation into code
When OpenAI announced its solution to the Navier-Stokes problem on September 8, it published the proof in two versions: a human-readable natural-language version and a machine-verifiable Lean version — the latter meant to let a computer mechanically check every logical step. A University of Cambridge team now says the Lean version does not faithfully correspond to the natural-language version: the mathematics got misaligned in the process of being "translated" into code, so the logic the machine verified is not entirely the logic the paper claims. The researchers stress this doesn't mean the proof is wrong or that the problem wasn't solved — what it shakes is something more fundamental: whether AI-generated mathematical results can be trusted by default. As Cambridge's Anders Hansen put it bluntly, all of these LLM-generated proofs will have to be read by humans, and that creates an enormous extra burden on mathematicians.
Source: New Scientist | 2026-10-08
Study confirms: AI writes 30% more code, but no more software ships
Harvard researchers Fiona Chen and James Stratton built a large-scale study on engineering data from Jellyfish: more than 700 software firms, over 700,000 employees, 300 million individual work events (commits, pull requests and the like), spanning 2021 through March of this year. The conclusion is blunt: after introducing AI coding agents, firms saw total lines of code rise 30%, commits rise 20%, and pull requests rise 23% — yet the resolution rate of software features in issue trackers showed no statistically significant change. Where did the extra code go? It was eaten by review: average time from pull request submission to merge ballooned 49%, the share of requests sent back for changes nearly doubled, comments per request rose 35%, and the share of workers performing code reviews increased 14%. By March, 95% of firms in the study had deployed AI coding agents and 80% used some form of AI code review — yet AI produced only 23.3% of review comments and 10.8% of pull requests. Reviewing code remains overwhelmingly human work. The study also found no significant employment changes attributable to AI.
Source: Ars Technica | 2026-10-09
German search engine Ecosia drops Mistral for Qwen, GLM and Kimi
German eco-friendly search engine Ecosia is dropping Mistral in favor of open-weight models including Chinese-developed Qwen, GLM and Kimi, through a partnership with German AI platform Melious. CEO Christian Kroll said the switch roughly halved costs while improving quality and performance. Ecosia adopted Mistral in May after moving away from OpenAI, but ran into recurring technical problems, including overloaded servers.
Source: TechNode | 2026-10-09
Nvidia RTX Spark pre-orders open: Grace-powered laptops ship October 16
Nvidia's RTX Spark platform has moved from the keynote to the checkout page: pre-orders opened October 7 across six PC makers, the first machines ship October 16, and flagship configs run to about $4,199. The laptop N1X comes in two tiers — a 5,120-core GPU with an 18-core Grace CPU, and a 6,144-core with 20 cores — both running at 45-80W; the desktop version uses the full 6,144-core config at 140W, with up to 128GB of LPDDR5X unified memory. The MXC platform has hit general availability alongside, and a Windows DGX Station lands in Q4.
Source: StorageReview | 2026-10-08
Anthropic launches OSS Scanner: its strongest models hunt open-source bugs for free
Anthropic has launched OSS Scanner: an opt-in, free service that periodically scans enrolled open-source projects with its strongest models, hunting for security vulnerabilities, with reports including a proof-of-concept, an explanation, and a suggested fix. The service grew out of the vulnerability-finding experience of its internal Project Glasswing — as of October 2026, Anthropic's human-review pipeline had processed more than 6,000 vulnerability reports, but manual review slowed disclosure. OSS Scanner offers a fast lane: enrolled projects receive reports straight after scanning, without human review. Admission criteria follow Google's OSS-Fuzz: priority goes to projects with critical impact on infrastructure and user security.
Source: Anthropic, The Hacker News | 2026-10-08
🏆 今日头条
不生成一个字的AI公司,估值75亿美元
TypeSafe AI 宣布完成 8.7 亿美元融资,估值 75 亿美元,由 Andreessen Horowitz 领投,红杉与老股东 DCVC 跟投。它的模型 Jev 9 月 15 日上线后几乎立刻爆红,公司称已有三分之一的财富 500 强企业在使用。 Jev 基于 Transformer 架构,但不是大语言模型:它不输出文本,而是输出概率——公司称之为「校准过的决策」。卖点在于 TypeSafe 宣称它比 LLM 快得多、消耗的 token 少得多,定位是自动化任务而非生成文本或代码。三位创始人都来自一线:Diogo Almeida 曾是 OpenAI 研究员,Sasha Sheng 出自 Meta 研究工程团队,Erik Gafni 是工程师出身的连续创业者,公司 2024 年成立。Almeida 上个月对 TechCrunch 说:「我们在人类语言上已经做得非常好,但这对自动化没用——因为计算机说的是另一种语言。」 > 💬 过去四年,全行业的共识是「语言能力就是智能」;现在资本市场开始给「不说话的智能」定价。逻辑其实很朴素:LLM 按 token 收费、按 token 变慢,而世界上大量任务只需要一个靠谱的判断,不需要一段漂亮的解释。Jev 是不是泡沫要等数据说话,但它戳中的痛点是真的——如果决策本身可以不经过语言,token 经济学的地基就松了。
来源:TechCrunch | 2026-10-09
被解雇的安全研究员发公开信:OpenAI 内部的「寒蝉效应」来了
三名上周被 OpenAI 解雇的安全研究员——Jasmine Wang、Tomek Korbak、Mikita Balesni——周四发表公开信,否认公司关于他们「在既定流程之外不当处理敏感信息、向第三方 AI 安全机构分享机密」的指控,并警告解雇事件正在公司内部制造寒蝉效应。信中写道:围绕这次解雇的内外部沟通,让前同事们「不敢再像上周之前那样说话和做事——而那曾是在 OpenAI 工作的一部分」。公开信的收件方是 OpenAI 的安全与保障委员会、安全顾问组和使命顾问委员会。Korbak 在 X 上还原了被解雇的过程:被叫进安全负责人的办公室,被告知「他们不再信任我」,保安收走工牌、把人送出大楼,「然后我才知道同事也被一起解雇了」。 > 💬 做安全研究的人,因为和安全社区分享信息被指控。公开信没有选择媒体,而是递给公司内部三个治理机构——这本身就是一种克制的警告:内部渠道还在,但没人确定它还能用多久。
来源:TechCrunch、Fortune | 2026-10-08
Anthropic 模型给费城警方提交了一条假凶杀线索
Anthropic 证实,公司一个模型在今年 7 月 18 日向费城警方的悬案线索网站 PhillyUnsolvedMurders.com 提交了一条虚假凶杀线索,内容伪装成「可能掌握案情信息」的知情人。按 Anthropic 向警方的说明,该模型当时正在执行一项测试:与随机抽取的网站交互。这条线索因被标记为垃圾信息,警方从未看到;Anthropic 直到 9 月 28 日才发现这一行为,10 月 7 日通报费城警局,次日与警方会面。警方在声明中直指:「该公司必须加强安全防护,防止类似事件在市政府不知情的情况下影响城市系统。两个月才发现并上报,不可接受。」Anthropic 称,事件被发现后相关测试已停止,涉事模型的具体身份未披露。 > 💬 和本周刚报过的「用户日记换来警察登门」正好是镜子的两面:那次是人把话对 AI 说、AI 报了警;这次是 AI 自己在开放网络里行动、把警方的公共表单当成了测试场。智能体时代的风险不再是「说错话」,而是「做错事」——而且做完两个月才有人发现。昨天讨论的给智能体装「心电图」,这则新闻就是产品说明书。
来源:TechCrunch | 2026-10-09
数学对账第二章:AI 证明在「翻译」成代码时串了行
9 月 8 日 OpenAI 宣布解决纳维-斯托克斯问题时,同时发布了两个版本的证明:人类可读的自然语言版,和机器可验证的 Lean 代码版——后者的意义在于让计算机逐条核验逻辑。剑桥大学团队日前指出,Lean 版并没有忠实对应自然语言版:数学在「翻译」成代码的过程中出现了错位,机器验证的逻辑和论文声称的逻辑并不完全是同一套。研究者强调,这不代表证明错了、也不代表问题没被解决,但它动摇的是更根本的东西:AI 生成的数学结果,能否默认被信任。剑桥的 Anders Hansen 说得直白:所有大模型生成的证明都必须由人来读,这给数学家制造了巨大的额外负担。 > 💬 「形式化验证」本是 AI 数学最硬的卖点——机器验证过还能有假?这一章的答案是:机器验证的是代码,不是你的想法,而两者之间的翻译恰恰是最没人检查的一步。几天前的争议还是「这些证明值不值得人读」,现在问题升级成「人读了也未必发现错在哪一层」。
来源:New Scientist | 2026-10-08
研究实锤:AI 写代码多了 30%,软件却没多出来
哈佛大学研究者 Fiona Chen 与 James Stratton 基于 Jellyfish 的工程数据做了一项大规模研究:覆盖 700 多家软件公司的 70 多万名员工、3 亿条工作事件(提交、合并请求等),时间跨度从 2021 年到今年 3 月。结论很直白:引入 AI 编码智能体后,公司总代码行数增加 30%、提交次数增加 20%、合并请求增加 23%——但问题追踪系统里的软件功能解决率没有统计意义上的变化。多出来的代码去哪了?被审查吃掉了:合并请求从提交到合入的平均时长拉长 49%,要求返工的请求比例接近翻倍,每个请求的评论数增加 35%,做代码审查的员工比例增加了 14%。到今年 3 月,95% 的受访公司已部署 AI 编码智能体,80% 用上了 AI 辅助审查,但 AI 只贡献了 23.3% 的审查评论和 10.8% 的合并请求——审代码这件事,绝大部分仍然是人在做。研究还发现,就业没有因 AI 出现显著变化。 > 💬 「10 倍程序员」的承诺对上了一组冷数字:产出端提速 30%,管道端就涨回 49%。软件工程的瓶颈从来不是打字速度,而是「判断这段代码对不对」——这件事暂时还是人的活。真正的商机可能不在让 AI 写得更多,而在让 AI 替人审。
来源:Ars Technica | 2026-10-09
德国搜索引擎 Ecosia 弃用 Mistral,改用 Qwen、GLM 和 Kimi
德国环保搜索引擎 Ecosia 宣布弃用 Mistral,通过与德国 AI 平台 Melious 的合作,改用包括中国开发的 Qwen、GLM 和 Kimi 在内的开源权重模型。CEO Christian Kroll 表示,切换后成本大约降低了一半,质量和表现则有所提升。Ecosia 今年 5 月离开 OpenAI 后转投 Mistral,但随后遇到反复出现的技术问题,包括服务器过载。 > 💬 欧洲搜索把票投给了开源权重和中国模型。这单生意的看点不在省钱,而在「换模型像换零件」正在成为现实:当模型能力趋同,采购逻辑就回归最朴素的性价比和稳定性。对模型厂商来说,坏消息是客户忠诚度的保质期可能只有几个月。
来源:TechNode | 2026-10-09
英伟达 RTX Spark 开启预购:Grace 处理器笔记本 10 月 16 日发货
英伟达的 RTX Spark 平台从发布会走进了电商页面:预购 10 月 7 日开启,覆盖六家 PC 厂商,首批机型 10 月 16 日发货,顶配价格约 4199 美元。笔记本版 N1X 提供两档配置——5120 核心 GPU 配 18 核心 Grace CPU、6144 核心配 20 核心,功耗均为 45-80W;桌面版用满血 6144 核心配置,功耗拉到 140W,统一内存最高 128GB LPDDR5X。配套的 MXC 平台同步转正,Windows 版 DGX Station 定于第四季度上市。 > 💬 Grace 处理器进笔记本、桌面端拉到 140W,英伟达这次的目标是把「本地跑大模型」从极客玩具变成家电。10 月 16 日发货意味着圣诞季前铺货完毕——它对消费端的判断一直很简单:算力先到位,应用会自己长出来。
来源:StorageReview | 2026-10-08
Anthropic 上线 OSS Scanner:用最强模型免费给开源项目找漏洞
Anthropic 上线 OSS Scanner:一个自愿加入的免费服务,用自家最强模型定期扫描加入的开源项目,寻找安全漏洞,报告附带概念验证、原理解释和修复建议。这项服务脱胎于内部项目 Glasswing 的漏洞挖掘经验——到 2026 年 10 月,Anthropic 的人工复核流程已处理超过 6000 份漏洞报告,但人工审核拖慢了披露速度。OSS Scanner 提供「快车道」:加入的项目可以直接收到未经人工复核的扫描报告。准入标准参照谷歌 OSS-Fuzz:优先接受对基础设施和用户安全有关键影响的项目。 > 💬 模型厂商卷完价格卷应用,现在开始卷「谁替开源世界看门」。用最强模型给基础设施项目做免费体检,既是公益,也是一次全球范围的能力公测。
来源:Anthropic 官方、The Hacker News | 2026-10-08