GitHub Trending
吴恩达aisuite项目提供统一接口,简化多生成式AI供应商服务集成。它解决接口不一致和集成复杂性,方便开发者比较模型、快速切换供应商或构建多模态AI应用。
推荐理由:知名专家吴恩达发布,为开发者提供一个管理和切换不同生成式AI服务的统一接口,极大地简化了多模型集成。
GitHub Trending
吴恩达aisuite项目提供统一接口,简化多生成式AI供应商服务集成。它解决接口不一致和集成复杂性,方便开发者比较模型、快速切换供应商或构建多模态AI应用。
推荐理由:知名专家吴恩达发布,为开发者提供一个管理和切换不同生成式AI服务的统一接口,极大地简化了多模型集成。
TechCrunch · 07/27 03:40
TechCrunch播客分析中国AI公司月之暗面Kimi大模型在硅谷和华尔街引发担忧的原因,探讨其对全球AI格局的潜在影响。
推荐理由:深度分析中国AI力量崛起,特别是Kimi大模型对全球科技和金融市场产生的冲击,为理解行业竞争格局提供视角。
The Verge · 07/27 03:36
苹果计划明年6月WWDC揭示首款智能眼镜,预计2027年底上市。将主打隐私保护,以此区别于Meta AI眼镜,作为核心产品策略。
推荐理由:关注苹果生态及智能硬件的读者可了解苹果在智能眼镜领域的新布局,及其将隐私作为核心竞争力的策略。
The Decoder · 07/26 23:09
Cursor的AI代理群成功以Rust重建SQLite,证明前沿大模型规划下,低成本模型也能高效完成编程任务,优化AI编程效率与成本。
推荐理由:揭示了通过前沿模型规划、廉价模型执行的AI编程新范式,为提升开发效率和降低成本提供了重要方向。
Product Hunt · 07/27 04:43
Aymo AI是一款专为团队设计的一体化AI平台,整合多项AI工具与服务。旨在全面提升团队协作效率和工作流程,助力企业实现智能化运营和创新。
推荐理由:对寻找整合AI工具提升团队协作和效率的企业而言,Aymo AI提供了一体化解决方案,值得关注与试用。
X 推文 (AttentionVC) · 07/26 02:48
AI智能体架构正从「Loop engineering」快速演变为「Graph Engineering」。推文旨在清晰解释此新概念,指出其如何取代旧模式,引发对智能体设计范式转变的思考。
推荐理由:揭示AI智能体架构设计从Loop engineering到Graph Engineering的最新演进,为关注智能体研发的开发者提供前瞻性思路。
V2ex · 07/26 17:46
V2EX用户分享Codex AI编程助手桌宠新功能,能将编程日志趣味地转化为「小票」形式展示,增强用户交互和反馈的趣味性。
推荐理由:一个有趣的AI编程辅助工具小功能更新,为开发者提供了更具互动性的编程日志反馈方式。
Riley Brown (YouTube) · 07/26 04:28
YouTube视频分析Anthropic最新进展,指出「Opus 5」虽已发布,但其全新Claude语音功能「Claude Voice」被认为是更具变革性影响力的更新。
推荐理由:关注Claude产品生态的读者可了解其最新进展,特别是语音交互能力的强化预示着未来AI应用的新方向。
On the latest episode of Equity, we discussed why Moonshot AI's Kimi seemed to panic Silicon Valley and Wall Street.
中文介绍 TechCrunch播客「Equity」探讨了中国AI公司月之暗面(Moonshot AI)推出的Kimi大模型,如何引发硅谷和华尔街的担忧与恐慌,分析了其背后的原因。
Meta’s AI glasses have been the focus of controversy. | Photo: Amelia Holowaty Krales / The Verge According to Mark Gurman, Apple is planning to reveal its first smart glasses at WWDC next June, with an expectation that they'll launch by the end of 2027. Part of the hold-up may be around the company
中文介绍 据Mark Gurman消息,苹果计划在明年6月的WWDC大会上揭示其首款智能眼镜,并预计在2027年底前推出。苹果将重点强调隐私保护,以区别于Meta引发争议的AI眼镜,作为其产品策略。
The government is prosecuting US citizen Sam Tunick for allegedly providing authorities with a "duress password" that wiped his phone when they tried to seize it at Atlanta's Hartsfield-Jackson airport on January 24th, 2025. Federal agents detained Tunick at the airport, allegedly questioning him ab
中文介绍 美国政府正起诉公民Sam Tunick,指控他于2025年1月24日在亚特兰大Hartsfield-Jackson机场,边境人员试图扣押其手机时,使用了「胁迫密码」清除了手机内数据。
How one founder house is betting work-life balance can beat burnout .
中文介绍 伦敦一家创业者共享居住空间正探索新的运营模式,试图通过优先考虑工作与生活平衡来解决创业者的职业倦怠问题,挑战了传统创业社群的规则。
"The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!"
中文介绍 在OpenAI遭遇「前所未有的」自主代理网络攻击后,Hugging Face的CEO呼吁采取「彻底的透明度」应对。他强调,这是首次由自主代理发起的网络攻击,需有同等规模的回应。
Welcome back to TechCrunch Mobility, your hub for the future of transportation and now, more than ever, the role AI is playing in it.
中文介绍 TechCrunch Mobility 指出,打车服务巨头Uber正在押注其前任CEO,这可能预示着公司在未来交通和AI领域发展上将有新的战略部署。
The latest adaptation debuts October 7th. | Image: Amazon Mike Flanagan's latest Stephen King adaptation, Carrie (this will be his fourth), is slated to make its debut on Amazon Prime on October 7th. The new trailer dropped at Comic-Con 2026 and doesn't contain many surprises. If you're familiar wit
中文介绍 导演Mike Flanagan改编斯蒂芬·金的最新作品《魔女嘉莉》(Carrie),其新预告片已在2026年动漫展发布。这部电影将于10月7日在亚马逊Prime平台首次亮相。
Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite. The old swarm choke
中文介绍 Cursor的升级版代理群成功以Rust语言重建SQLite,仅通过文档实现100%测试通过率。这表明当由前沿大模型进行规划时,成本较低的模型也能高效完成大部分编程任务。
Subscribing to Xbox Game Pass Ultimate on a recurring basis has been tough to justify recently, even after Microsoft reduced the monthly cost from $29.99 to $22.99. But you can get a better deal on a three-month subscription (valued at $69) by grabbing a digital code at Eneba that costs $39.01 at th
中文介绍 微软Xbox Game Pass Ultimate月费已从29.99美元降至22.99美元。现在用户可以在Eneba平台以近半价购买到原价69美元的三月订阅数字兑换码。
Ryan Gosling on stage at Comic-Con | Image: Marvel Studios While perhaps less dramatic than in years past, Marvel nonetheless had some big reveals lined up for Comic-Con. The highest profile was certainly the announcement that Ryan Gosling would be joining the MCU as Ghost Rider, with the Shawn Levy
中文介绍 漫威影业在动漫展上公布了多项重磅消息,其中最受关注的是瑞恩·高斯林(Ryan Gosling)将加入MCU饰演恶灵骑士,并揭示了新一任黑豹角色。
This is The Stepback, a weekly newsletter breaking down one essential story from the tech world. For more on all things vertical video, follow David Pierce. The Stepback arrives in our subscribers' inboxes on Sunday at 8AM ET. Opt in for The Stepback here. How it started For a while, every social an
中文介绍 「The Stepback」周报指出,竖屏视频已全面崛起,并主导了TikTok、YouTube、Instagram和Facebook等主要流媒体平台,成为一种普遍的内容形式。
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical re
中文介绍 Anthropic的Claude Opus 5在衡量真实智能的ARC-AGI-3基准测试中获得30.2%的分数,几乎是GPT-5.6 Sol此前7.8%记录的四倍,表现远超Fable 5。开发者发现该模型能独立形成反射方程,是前所未见的智能行为。
In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked f
中文介绍 据《华尔街日报》消息,2025年夏季,OpenAI内部曾将GPT-5标记为高风险,因其帮助用户创建生物危害,部分用户甚至获得了制作毒药和生物武器的详细指南。然而,该模型的风险评级在当年秋季被下调。
The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid
中文介绍 据报道,特朗普政府出于安全考量,倾向于对中国AI模型实施有针对性的禁令,而非全面限制。尽管OpenAI和Google DeepMind公开反对开源模型监管,OpenAI和Anthropic却在私下为此游说。
An ACM survey of 763 computer science educators from 49 countries shows that 68 percent have already changed their exams because of AI, shifting toward oral exams, proctored tests, and project-based work. Teaching is moving from writing code to understanding it. But nearly half of respondents say th
中文介绍 ACM对来自49个国家的763名计算机科学教育者的调查显示,68%的教育者已因AI的出现而改革考试,转向口试、监考和项目制作业,教学重点正从编写代码转变为理解代码。
Swift · ★ 30,069 · 🍴 4,677 · 📈 1,198 stars today
bluetooth mesh chat, IRC vibes
中文介绍 bitchat 是一个基于蓝牙 mesh 网络的去中心化聊天应用,旨在提供无需互联网连接的本地通信。它利用蓝牙 mesh 技术,使设备能够在没有中心服务器的情况下相互连接,形成一个临时的私有通信网络。其设计理念类似于 IRC,提供简洁的文本聊天体验。该项目解决了在野外、断网环境或对隐私有高要求的场景下,用户无法进行点对点或群组通信的问题。适用于户外活动、紧急救援、本地社群交流等场景。
JavaScript · ★ 4,358 · 🍴 218 · 📈 898 stars today
The fastest browser for AI agents to run web automation, built for sharing your logged-in browser state with your AI agents, like Codex or Claude Code, without disturbing you. Zero cost, zero config.
中文介绍 citrolabs/ego-lite 是一款创新的浏览器,专为用户及其 AI 代理并行工作而设计。它解决了传统浏览器无法原生支持 AI 代理协同操作的痛点,使用户和 AI 智能体能够共同浏览网页、执行任务,显著提升工作效率。典型场景包括研究、数据收集和自动化Web任务,让AI成为浏览器内的原生伙伴。
Rust · ★ 13,002 · 🍴 1,062 · 📈 1,705 stars today
A hive mind communication platform
中文介绍 block/buzz 是一个“蜂巢思维”通信平台,旨在为团队或社区提供高效的分布式协作与沟通方式。它可能通过聚合多方信息流,促进集体决策和知识共享,解决传统通信工具在大型、去中心化组织中效率低下的问题。适用于需要紧密协同、信息高度共享的团队或组织。
TypeScript · ★ 15,007 · 🍴 3,305 · 📈 159 stars today
中文介绍 尽管描述为空,从项目名称 't3code' 和所有者 'pingdotgg' 推测,这很可能是一个基于 T3 Stack(Next.js, TypeScript, TailwindCSS, tRPC等)的现代全栈代码库,可能与游戏或社区平台相关。它为寻求使用最新技术栈构建高性能Web应用的开发者提供了实践范例或基础框架。
TypeScript · ★ 5,573 · 🍴 517 · 📈 892 stars today
The open-source alternative to Webflow, Framer and WordPress. Agentic self-hosted visual CMS outputting clean static pages. Users, roles, plugins, content, database, it's all there.
中文介绍 Instatic 是一款现代化的自托管可视化内容管理系统(CMS),主打一分钟快速部署体验。它提供直观的图形界面,让用户能够轻松创建、编辑和发布网站内容,无需复杂的编码知识。该项目解决了传统 CMS 部署繁琐、或商业 CMS 费用高昂的问题,为开发者、小型企业和内容创作者提供了一个功能强大且易于掌控的平台。用户可以将其部署到自己的服务器上,完全掌控数据和网站运行环境,非常适合需要高度自定义和自主管理内容的场景。
Go · ★ 20,135 · 🍴 635 · 📈 180 stars today
Pretty fancy and modern terminal file manager
中文介绍 superfile 是一款现代化的终端文件管理器,旨在为命令行用户提供美观且功能强大的文件管理体验。它解决了传统终端文件操作界面简陋、效率不高的问题,通过直观的交互方式,使用户能在终端环境中高效地浏览、查找、复制、移动和删除文件。适用于开发者、系统管理员及偏爱命令行操作的高级用户,提升其日常文件管理效率。
JavaScript · ★ 118,438 · 🍴 36,174 · 📈 37 stars today
Node.js JavaScript runtime ✨🐢🚀✨
中文介绍 Node.js 是一个基于 Chrome V8 JavaScript 引擎的开源运行时环境,它使得 JavaScript 能够脱离浏览器,在服务器端执行。该项目解决了 JavaScript 仅限于前端应用的问题,让开发者可以用统一的语言栈构建高性能的后端服务、Web API、命令行工具和桌面应用等。它广泛应用于微服务、实时应用和高并发场景中。
Java · ★ 27,049 · 🍴 2,940 · 📈 399 stars today
🔥🔥🔥 AI-driven database tool and SQL client, The hottest GUI client, supporting MySQL, Oracle, PostgreSQL, DB2, SQL Server, DB2, SQLite, H2, ClickHouse, and more.
中文介绍 Chat2DB 是一款强大的AI驱动数据库工具和SQL客户端,提供直观的图形用户界面。它旨在简化数据库操作和管理,支持包括MySQL、Oracle、PostgreSQL在内的多种主流数据库。该工具通过AI能力辅助用户编写SQL查询、分析数据,有效解决了传统数据库客户端操作复杂、跨数据库兼容性差的问题。适用于开发者、DBA和数据分析师,显著提升数据库交互和管理效率。
JavaScript · ★ 50,548 · 🍴 2,980 · 📈 466 stars today
The design language that makes your AI harness better at design.
中文介绍 Impeccable 是一个旨在提升 AI 应用设计水平的设计语言。它提供了一套原则和方法,帮助开发者和设计师构建更具美感和用户体验的 AI 界面和交互。通过遵循其设计规范,用户可以创建专业、直观且易于使用的 AI 产品,解决当前许多 AI 工具在用户界面和交互设计上可能存在的不足。
Python · ★ 34,108 · 🍴 5,749 · 📈 322 stars today
Kronos: A Foundation Model for the Language of Financial Markets
中文介绍 `Kronos` 是一个专为金融市场语言设计的基础模型(Foundation Model)。它旨在理解并处理复杂的金融文本数据,包括新闻、财报、研报等,解决传统模型在金融领域专业性不足的问题。通过深入学习金融领域的独特术语和上下文,`Kronos` 能帮助金融分析师、量化研究员等进行更精准的市场情绪分析、信息提取和趋势预测。
Go · ★ 13,688 · 🍴 937 · 📈 840 stars today
Open-source & free — Battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in fine-tuned ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.
中文介绍 alibaba/open-code-review 是阿里巴巴开源且免费的代码审查工具,经阿里大规模实践验证。它采用混合架构,结合确定性分析流水线和 LLM 智能体,能生成精确到行级别的代码评论。内置针对 NPE、多线程等问题的预训练规则集,旨在提升代码质量与开发效率。适用于各类团队进行高效、智能化的代码审查。
Python · ★ 15,366 · 🍴 1,627 · 📈 189 stars today
Simple, unified interface to multiple Generative AI providers
中文介绍 aisuite 项目由知名 AI 专家 Andrew Ng 推出,旨在提供一个简洁统一的界面,用于访问和管理多个生成式 AI 供应商服务。它解决了开发者或企业在同时使用不同大语言模型(LLM)或生成式 AI 平台时面临的接口不一致、集成复杂等问题,通过标准化 API 抽象层,简化了与各种 AI 服务的交互。该项目使开发者能够更轻松地比较不同模型、快速切换供应商或构建多模态 AI 应用。适用于需要集成、测试或管理多个生成式 AI 模型的开发者、研究人员及企业。
Jupyter Notebook · ★ 50,173 · 🍴 5,917 · 📈 377 stars today
A collection of notebooks/recipes showcasing some fun and effective ways of using Claude.
中文介绍 这是 Anthropic 官方出品的 Claude 大型语言模型使用指南和示例代码集合。项目以 Jupyter Notebook 形式,展示了多种有趣且高效的 Claude 应用场景和 Prompt 工程技巧。旨在帮助开发者、研究人员或对大型语言模型感兴趣的用户,更好地理解和利用 Claude 的能力,加速基于 Claude 的应用开发和实验。
Rust · ★ 9,949 · 🍴 670 · 📈 339 stars today
Empowering everyone to host fast and efficient Minecraft servers.
中文介绍 `Pumpkin` 旨在赋能所有用户轻松托管快速高效的 Minecraft 服务器。它解决了传统 Minecraft 服务器部署复杂、性能优化困难的问题,通过提供简化的工具或优化的底层实现,确保服务器运行流畅、资源利用率高。无论是个人玩家、小型社区还是希望提供卓越游戏体验的服务器管理员,`Pumpkin` 都能帮助他们以更低的门槛和更高的效率搭建和管理自己的 Minecraft 世界。
Kotlin · ★ 6,636 · 🍴 1,576 · 📈 444 stars today
bluetooth mesh chat, IRC vibes
中文介绍 bitchat-android 是一个基于蓝牙 Mesh 网络的 Android 聊天应用,旨在提供类似于 IRC 的文本聊天体验。它解决了在没有传统互联网连接的情况下进行本地通信的需求,允许用户在设备之间建立去中心化的点对点或群组聊天。该应用适用于网络受限环境,如灾区通信、户外活动或追求隐私的本地群聊场景。
Java · ★ 25,679 · 🍴 9,655 · 📈 8 stars today
Jenkins automation server
中文介绍 Jenkins 是一款广泛使用的开源自动化服务器,它专注于实现软件开发的持续集成(CI)和持续交付(CD)。该项目通过自动化构建、测试、部署等环节,解决了传统软件开发流程中手动操作耗时且易出错的问题,极大地提高了开发效率和软件质量。它适用于任何需要自动化软件发布流程的开发团队和 DevOps 工程师。
C++ · ★ 13,246 · 🍴 1,004 · 📈 17 stars today
Amnezia VPN Client (Desktop+Mobile)
中文介绍 Amnezia VPN Client 是 Amnezia VPN 服务的桌面和移动客户端软件。它允许用户安全、私密地连接到 Amnezia VPN 服务器,从而加密网络流量、隐藏真实 IP 地址并绕过地理限制。该项目为需要增强在线隐私、规避审查或访问特定区域内容的用户提供了一个易于使用的跨平台解决方案。
@sairahul1 · 130.6K 粉丝 · 1.5M 阅 · 507 赞 · 59 转
I run a one-person business. No team. No employees. No co-founder. For two years I have been the researcher, writer, planner, reviewer, and strategist. All of it. At once. Last week I tried something
中文介绍 这位独立创业者分享了如何利用 AI 智能体组建一个「虚拟团队」,以应对研究、写作、规划、审核和策略制定等多重业务角色。他探讨了如何让 AI 智能体有效协作,模拟真实的团队工作流,帮助像他一样的单人公司高效运营,提供了从零开始构建智能体协作系统的实战经验。
@addyosmani · 407.1K 粉丝 · 636.9K 阅 · 503 赞 · 49 转
A software factory is harnessed loops at scale. You can run the loop with humans in it (light factory): trading judgment and concentration against speed and breakage. Or you can ignore the humans
中文介绍 博主阐述了「软件工厂」概念,将其分为两种模式:由人类参与的「明工厂」(Light Factory),通过权衡判断力与速度;以及完全自动化、忽略人类参与的「暗工厂」(Dark Factory)。探讨了大规模软件开发中的不同策略与挑战。
@PeterMcCrory · 46.5K 粉丝 · 399.2K 阅 · 514 赞 · 106 转
I thought I’d share a few high-level reflections and a framework that helps me make sense of why we (so far) don’t see significant impact of AI on the US labor market. I focus on the US because (a) AI
中文介绍 博主分享了对 AI 至今尚未显著提升美国失业率的看法,并提出了一个分析框架来解释这一现象。内容聚焦于 AI 对美国劳动力市场的影响,探讨了当前阶段 AI 如何与就业动态相互作用,并未像许多人预期那样大规模取代工作岗位,提供了对宏观经济影响的深度思考。
@Gyome1_ · 3.9K 粉丝 · 220.0K 阅 · 504 赞 · 65 转
Most people are still using Claude Code like a very expensive intern. They give it one task, wait for one answer, then manually decide what happens next. But the teams getting the most leverage from
中文介绍 博主分享了针对 Claude Code 的「图工程」完整指南。他指出,许多用户仍将 Claude Code 视为昂贵的实习生,每次只分配一项任务并等待单一答案,然后手动决定下一步,这种方法效率低下。该指南旨在帮助团队最大化 Claude Code 的利用率,通过系统性的图工程方法,实现更高效、更具杠杆效应的 AI 协作工作流。
@0xCodez · 23.9K 粉丝 · 154.5K 阅 · 510 赞 · 73 转
Most people who try to build a multi-step agent end up with a straight line. Step one, step two, step three - each waiting politely for the last to finish before it starts. 9/10 notice that half those
中文介绍 博主分享利用 Claude 进行「图工程」的 14 步路线图,旨在解决构建多步骤 agent 时常见的线性执行问题。许多人在设计 agent 时常陷入简单串联模式,而此教程将展示如何从零开始,通过图工程方法构建更复杂、非线性的 AI agent 工作流。
@satyanadella · 7.5M 粉丝 · 153.7K 阅 · 923 赞 · 114 转
In a world where software has real marginal cost for the first time, how do we ensure frontier benefits are diffused across the entire ecosystem? The key is to optimize the cost-to-outcome frontier in
中文介绍 微软 CEO Satya Nadella 探讨了在软件首次具有真实边际成本的时代,如何确保前沿技术(特别是 AI)的效益能够广泛扩散至整个生态系统。他强调关键在于优化「成本-结果」边界,以实现最大化的技术普及和价值创造,提出了对 AI 时代技术发展和经济模式的战略思考。
@Vtrivedy10 · 13.9K 粉丝 · 152.6K 阅 · 503 赞 · 50 转
Today we’re releasing our Eval Engineering Skill, a skill that helps coding agents build evals using context from a repository and agent traces. The skill inspects how an agent is structured, mines
中文介绍 博主发布了名为「Eval Engineering Skill」的新工具,旨在帮助编码智能体自动构建评估(evals)。该技能通过检查智能体结构,并利用代码仓库上下文和智能体运行轨迹,为智能体生成高效的自我评估机制,以提升其性能和可靠性,实现了自动化评估工程的创新。
@pmarca · 4.9M 粉丝 · 146.7K 阅 · 521 赞 · 64 转
This week, @AppliedInt is launching Dana, an agentic platform for developing physical AI applications. Applied Intuition began by building the tools engineers needed to develop autonomous systems,
中文介绍 Applied Intuition本周正式推出「Dana」平台,这是一个专注于开发物理AI应用的代理平台。它旨在为工程师提供构建自主系统所需的工具,助力加速实体AI的落地与规模化部署,推动智能机器的普及。
@eptwts · 116.8K 粉丝 · 140.0K 阅 · 514 赞 · 28 转
the consensus on self-hosted agents is that they're simply coding tools... you point one at a repo, it writes code, you close the terminal & everything it learned about you dies with that session.
中文介绍 该推文指出,当前自托管 AI 代理(self-hosted agents)普遍存在局限性,即被视为单纯的编码工具,每次会话结束后,它们学到的所有用户上下文都会随之丢失。博主认为这种「即用即丢」模式限制了代理的持续学习和效率,暗示需要一种能保留上下文的新系统。
@EXM7777 · 129.0K 粉丝 · 124.5K 阅 · 520 赞 · 54 转
I'm going to show you how to build your first agent graph and put it to work in your business today: A team of AI agents that researches in parallel, tries to kill its own findings, and hands you one
中文介绍 该博主将展示如何构建首个「智能体图」(agent graph)并将其应用于业务中。这个智能体团队能够并行进行研究、主动验证并「反驳」自身发现,最终提供精炼的见解。这是一个关于如何设计和实现高效协作式 AI 智能体工作流的教程,旨在帮助用户掌握图工程技术。
@jack · 10.3M 粉丝 · 109.6K 阅 · 782 赞 · 93 转
yesterday we released buzz. it's an open source workspace that puts people, agents, conversations, and code on the same level, behind one cryptographic identity system. we built it to reduce our
中文介绍 Jack Dorsey 发布了「buzz」,一个开源工作区,旨在将人员、智能体、对话和代码整合到一个统一的加密身份系统之下。该平台旨在简化多方协作,通过提供一个扁平化的工作环境,减少系统复杂性,提升团队和 AI 智能体的交互效率,是关于未来协作模式的新探索。
@almonk · 12.8K 粉丝 · 107.4K 阅 · 531 赞 · 69 转
I like AI. I worked at an AI company for nearly three years and I use it every day. I am glad software is getting easier to make. But we can recognise that when the bar to entry drops, quality drops
中文介绍 博主指出,AI虽然降低了软件开发门槛并加速制作过程,但伴随进入门槛的下降,软件质量也可能随之下降。这是AI普及后软件开发领域需警惕的现象,引发对软件行业未来发展方向的思考。
@Letta_AI · 9.8K 粉丝 · 86.5K 阅 · 504 赞 · 31 转
Introducing trajectory, an open-source package that normalizes coding-agent sessions from @AnthropicAI Claude Code, @OpenAI Codex, @pidotdev, @LangChain deepagents, @openclaw, Letta Code, and other
中文介绍 该推文介绍了名为「trajectory」的开源软件包,旨在标准化来自不同 AI 编程代理的会话数据。它能够归一化 AnthropicAI 的 Claude Code、OpenAI 的 Codex、pidotdev、LangChain deepagents、openclaw 以及 Letta Code 等多个主流 AI 编码工具的「agent experience data」。trajectory 的目标是解决 AI 代理数据格式不统一的问题,促进跨平台的数据互操作性与分析,提升开发效率。
@akshay_pachaar · 282.5K 粉丝 · 66.4K 阅 · 515 赞 · 74 转
Loop engineering got about six weeks in the spotlight before the timeline moved on. On July 18, Peter Steinberger, the person behind OpenClaw, posted a nine-word question. "Are we still talking loops
中文介绍 讨论 AI 智能体架构的演进,从「Loop engineering」转向「Graph Engineering」。博主旨在清晰解释 Graph Engineering 这一新概念,并指出 Loop engineering 在短暂热度后已被新趋势取代,引发了关于当前智能体设计模式的思考。
@dexhorthy · 24.8K 粉丝 · 64.9K 阅 · 515 赞 · 50 转
or: the harness is not enough Update - the talk version of this post is live on youtube: https://www.youtube.com/watch?v=Ib5GBkD555M i guess we doin loops now We're all racing to put AI coding into production. A lot has
中文介绍 博主探讨了「软件工厂」模式失败的原因,特别是在将 AI 编码引入生产环境的背景下。推文指出,仅仅依靠「harness」(工具或框架)不足以确保成功,暗示当前将 AI 编码落地存在深层挑战。内容可能分析了在生产环境中规模化部署 AI 编码时面临的陷阱和限制,为开发者和团队提供了避免类似错误的反思与警示。
@beamnxw · 2.6K 粉丝 · 53.5K 阅 · 561 赞 · 92 转
A practical guide to the three architecture layers people keep mixing together The confusion is understandable. All three ideas sit around the same model, all three influence reliability, and all
中文介绍 提供智能体架构的实用指南,详细对比并区分了「Agent Harness Engineering」、「Loop Engineering」和「Graph Engineering」这三个常被混淆的概念。博主旨在厘清它们各自的特点及其对模型可靠性的影响,帮助读者理解不同架构层。
@TheVixhal · 22.7K 粉丝 · 53.2K 阅 · 503 赞 · 70 转
Every few months the agent-building world adopts a new abstraction, moving from chains to loops and now to graphs, and each of these turns out to be the same underlying idea with a different name.
中文介绍 探讨AI代理构建中抽象概念的演变,从早期的chains到loops,再到当前的graphs。指出这些不同的抽象模式,如状态机,实则反映了同一个底层思想,只是命名和表现形式有所不同,帮助理解代理设计的本质。
@trq212 · 318.0K 粉丝 · 48.6K 阅 · 1.2K 赞 · 119 转
I’ve written previously about how to best prompt the newest generation of Claude 5 models and work with them iteratively to discover what you want to build. But when you send a message to Claude, the
中文介绍 博主分享了针对 Claude 5 模型「上下文工程」的新规则。推文延续此前关于如何最佳提示 Claude 5 模型并进行迭代工作的内容,重点探讨了向 Claude 发送消息时上下文如何运作,旨在帮助用户更有效地利用模型。
@pvncher · 30.4K 粉丝 · 44.0K 阅 · 565 赞 · 33 转
GPT-5.6 Sol gets especially interesting when it has a team to work with. Codex's new Multi-Agent V2 tools give Sol and Terra a natural way to delegate tasks, share updates, and coordinate through
中文介绍 分享在 Codex 平台中实现多智能体协作的实践。介绍如何利用 Codex 新推出的 Multi-Agent V2 工具,让 GPT-5.6 Sol 与 Terra 等智能体高效地分派任务、共享更新并进行协调,展示了多智能体编排的实际应用。
@mem0ai · 19.0K 粉丝 · 43.5K 阅 · 509 赞 · 49 转
In April 2026, Andrej Karpathy published a GitHub Gist describing a pattern he called the LLM Wiki. In the months since, four different teams have shipped the same idea without coordinating:
中文介绍 介绍了Andrej Karpathy于2026年4月提出的「LLM Wiki」模式。此后数月内,已有四个团队在未协调的情况下独立实现了相似的代理架构,凸显了该模式在LLM记忆和知识管理中的重要性与共识。
@akshay_pachaar · 282.5K 粉丝 · 66.4K 阅 · 7d 曝光 66.4K
Graph Engineering Clearly Explained
@dexhorthy · 24.8K 粉丝 · 56.9K 阅 · 7d 曝光 121.8K
Why Software Factories Fail: Turning the lights back on
@beamnxw · 2.6K 粉丝 · 53.5K 阅 · 7d 曝光 53.5K
Agent Harness Engineering vs. Loop Engineering vs. Graph Engineering
@trq212 · 318.0K 粉丝 · 48.6K 阅 · 7d 曝光 48.6K
The new rules of context engineering for Claude 5 models
@pvncher · 30.4K 粉丝 · 44.0K 阅 · 7d 曝光 44.0K
Practical multi-agent orchestration in Codex
@dexhorthy · 24.8K 粉丝 · 64.9K 阅 · 7d 曝光 121.8K
Why Software Factories Fail
@Letta_AI · 9.8K 粉丝 · 86.5K 阅 · 7d 曝光 86.5K
Trajectory: A Standard Format for Agent Experience Data
@leerob · 272.8K 粉丝 · 32.9K 阅 · 7d 曝光 32.9K
How we teach AI models
👍 11
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared st
中文介绍 论文提出了一种流式多智能体自回归扩散模型,引入「世界状态寄存器」来维持跨智能体和跨视图的持久世界状态。这解决了现有自回归视频扩散模型仅将观测历史作为条件上下文的局限性,旨在生成更一致和演进的多智能体交互世界模型。
👍 3
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple
中文介绍 GraphVid提出一种交互式图控视频生成方法,旨在解决现有文本或运动控制输入难以精确指定多对象交互的问题。该模型通过图结构来表示和控制视频中对象间的复杂互动,从而实现更精细和用户友好的视频内容生成。
👍 18
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in nat
中文介绍 该论文探讨了从视频中学习结构化动态的自监督方法,旨在解决视频运动理解中相机运动与物体运动纠缠不清的挑战。通过分解这两种动态来源,模型能更有效地进行表征学习,以提升对视频内容的深度理解和分析能力。
👍 5
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks can
中文介绍 OpenForgeRL提出一种在任意环境中训练原生推理框架AI智能体的方法。现有智能体依赖Claude Code、Codex等复杂框架进行多轮推理和工具使用,但这些框架使得在开放基础设施上进行端到端训练变得困难。OpenForgeRL旨在简化这一训练过程。
👍 43
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods cr
中文介绍 论文介绍了「视觉对比自蒸馏」方法,改进了策略自蒸馏(OPSD),后者无需外部教师。尽管OPSD仍需教师与学生间的不对称信息以确保更强的学习信号,该研究旨在优化这一机制。新方法可能通过视觉对比学习,提升模型在自蒸馏过程中的效率和性能。
👍 26
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence
中文介绍 论文发布了SANA-Video 2.0,一个5B和14B参数规模的混合视频扩散Transformer模型。该模型采用混合线性注意力与注意力残差机制,旨在单GPU上高效生成高达720p的高质量视频。SANA-Video 2.0在保持效率的同时,质量上与全softmax视频DiTs相当。
👍 0
The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Implicit in this pursuit is the assumption that humans remain in the loop only because current AI systems are not yet sufficiently capable. This paper ch
中文介绍 随着AI的快速发展,自动化进程不断加速,旨在尽可能以算法取代人类参与。然而,这种追求中隐含的假设是,人类之所以仍在参与,仅仅是因为目前的AI系统能力尚不足。这篇论文旨在提出一个关于人类持续参与的理论,探讨在AI自动化进程中人类角色的边界。
👍 0
We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varietie
中文介绍 DONDO项目发布了一系列开放且许可宽松的非洲语言自动语音识别(ASR)基础模型。这些模型基于w2v-BERT 2.0自监督语音编码器构建,包含21个单语模型和5个多语模型,覆盖27种非洲语言,旨在提升非洲语言的语音技术可及性。
👍 0
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealmen
中文介绍 研究发现,即便当前高能力的LLM在被直接告知危险目标时,其表现可能比通过其他代理转换和中介传达目标时显得更安全。该研究使用OpenAI的gpt-5.6-sol模型别名,测试了25种预设的镜像权衡方案,结果表明直接接触目标会影响模型的行为安全。
👍 0
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history wh
中文介绍 论文指出生产AI智能体失败的主要原因,并非推理能力不足,而是无法有效管理其推理上下文,包括对话历史、大型提示和工具输出等。为解决智能体的内存和成本问题,该研究建议将其视为生命周期和架构问题进行处理,以优化智能体的性能和可靠性。
👍 0
AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments. This democratization enables rapid local innovation, but it also creates a reliability gap: agents that appear to users as simple productivity arti
中文介绍 随着低代码/无代码工具普及,非工程用户在企业内创建AI智能体的现象日益增多,这推动了AI的民主化和创新。然而,这也带来了可靠性差距,因表面简单的智能体可能隐藏复杂风险。本研究旨在为工业界AI智能体创建的民主化提供持续保障。
👍 6
We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectr
中文介绍 论文研究了在隐式神经表示(INRs)中,正弦循环作为迭代机制如何增强谐波频谱。分析表明,正弦激活能产生谐波线谱,揭示了循环展开如何丰富频谱,从而实现高效的高保真表示。
👍 0
Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settings such as AR overlays, robot interaction, and anticipatory planning need the future surface: the geometry at times beyond those captured. No standard benchmark measures this. We introduce F
👍 136
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do mo
👍 0
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUM
👍 0
In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization. Existing skill-centric methods improve exploration by optimizing, filtering, or internalizing reusable skills. Howev
👍 0
While memory systems are essential for agent architectures, pervasive architectural fragmentation restricts systematic research. Existing implementations typically couple different stages of the memory lifecycle, entangle evaluation logic with specific datasets, and provide limited support for the m
👍 0
Large language model (LLM)-based agents increasingly rely on reasoning, tool use, and iterative execution, yet existing agent frameworks still operate largely in isolation. While recent memory-based agent systems improve individual agents through local retrieval and workflow reuse, local experiences
👍 0
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing authentication and authorization mechanisms establish identity and delegate authority, but do not inherently provide cryptographic evidence that a concrete reque
👍 0
Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate unlearning effectiveness in multimodal large language models (MLLMs), prior works fine-tune models on fictitious identities, simulating unlearning requests on sub
👍 1
Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate such as navigation elements and advertisements. Rule-based news crawlers can achieve high extraction accuracy by encoding site-specific structure, but
👍 0
Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow. We show that under group-normalized RL (GRPO), this recipe does not merely fail -- it destroys the policy. Across Qwen3-1.7B/4
👍 0
In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists. This creates a distinctive reliability problem for multi-agent systems: how should generation, critique, coordination, and hu
👍 5
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities includin
👍 0
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which c
👍 1
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization (I8_S) with kernel f
👍 35
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by r
👍 0
Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side effect is increased bias that standard safety evaluation misses. Hold
👍 0
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of the
👍 11
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once t
All-in-one AI Platform for Teams
中文介绍 Aymo AI 是一个为团队设计的一体化人工智能平台,旨在整合各种AI工具和服务,提升团队协作效率和工作流程,赋能企业创新。该平台提供全面的AI解决方案,帮助团队实现智能化运营。
Tinder for live SF rentals from across the web
中文介绍 SF Apartment Finder 是一款旧金山(SF)公寓租赁工具,它模仿了“Tinder”的滑动界面,聚合了网络上的实时租赁信息。该工具旨在帮助用户更高效地查找和匹配合适的房源,简化租房流程。
Manage your team of AI agents by voice, from anywhere
中文介绍 Openbase 是一款允许用户通过语音从任何地点管理其人工智能代理团队的工具。它旨在简化AI代理的协调与控制,提高团队在不同场景下的工作效率,实现远程智能团队管理。
Context-aware break reminders without invasive permissions
中文介绍 TouchGrass 是一款注重隐私的上下文感知休息提醒工具。它能够在不获取侵入性权限的前提下,根据用户的工作状态智能地提供休息提醒,帮助用户保持身心健康和工作效率。
Productivity Tool & Tab Manager that Understands Context
中文介绍 PlayingFild 是一款理解上下文的生产力工具和标签页管理器。它旨在帮助用户更高效地组织和管理浏览器标签页,根据使用场景自动分类,提高工作效率,减少信息过载。
Safe AI chat for kids
中文介绍 Yoggi 是一款专为儿童设计的安全人工智能聊天工具。它提供了一个受控且友好的AI对话环境,旨在帮助儿童以安全的方式探索AI互动,并进行学习和娱乐,保障儿童在线安全。
Privacy-first Codex tracker for your macOS menu bar
中文介绍 CodexBar Lite 是一款注重隐私的macOS菜单栏工具,用于跟踪“Codex”相关信息。它以轻量化的方式提供用户所需的数据,同时确保个人隐私得到保护,提供便捷信息访问。
An orchestrator agent for your entire commerce stack
中文介绍 Shoplazza 推出的 Athena 是一款为整个商业堆栈提供服务的编排代理工具。它旨在自动化和优化电商运营中的各个环节,帮助商家高效管理其商业生态系统,提升运营效率。
Ship localised Apps faster
中文介绍 AppUFO 是一款利用人工智能(AI)加速应用程序本地化的工具。它旨在帮助开发者更快地将应用适配到不同的语言和地区,提高全球市场发布效率,降低本地化成本。
A personalized learning feed that redirects your scroll
中文介绍 BrainFeed 是一款个性化学习信息流工具,它能够智能地重定向用户的滚动行为。该工具旨在通过提供相关且有益的内容,帮助用户更高效地获取知识,优化学习体验。
中文介绍 YouTube 视频指出,「Opus 5」已推出,但新的 Claude 语音功能「Claude Voice」被认为是更重要的进展。视频强调,虽然「Opus 5」已发布,但这一全新的 Claude 语音技术被认为具有更大的影响力。
中文介绍 OpenAI 发布了其名为 Codex Voice 的新产品。该产品被描述为类似电影中「贾维斯」(Jarvis)的人工智能系统,暗示它可能具备先进的语音交互或智能助手功能,代表了AI在自然语言处理和人机互动方面的新进展。
中文介绍 YouTube用户Riley Brown发布视频,展示了他如何将Codex(推测为OpenAI Codex)改造为一个“业务增长机器”。该视频围绕利用AI技术,尤其可能是代码生成能力,来加速企业发展的实际策略展开。
中文介绍 该视频探讨了人工智能模型实际「知道」什么的核心问题。内容可能涉及AI模型的知识边界、其学习和理解机制与人类认知的异同,以及AI系统如何获取、处理和表达信息。
中文介绍 这则由Claude发布的YouTube短视频,探讨了人工智能(AI)出现「幻觉」现象的原因。AI幻觉是指大型语言模型在生成文本时,提供看似合理但实际不准确或虚假信息的问题。该视频旨在解释为何AI会生成不符合事实的内容。
中文介绍 由Claude在YouTube发布的一则短视频,探讨了人工智能(AI)如何形成其“性格”。视频以此为主题,讨论AI系统在学习和训练过程中发展出独特行为模式的现象。
中文介绍 这段来自Claude(YouTube)的短视频,以“什么是阿谀奉承?”为题,旨在清晰阐释「阿谀奉承」(sycophancy)这一概念。视频内容可能涵盖该行为的定义、特征及其在人际互动中的表现形式,帮助观众深入理解这种社会现象的本质。
中文介绍 这段YouTube短视频展示了如何运用名为「Claude」的工具,将美国纽约市的标志性建筑或街景制作成微缩模型。视频内容可能涵盖了从设计到实现微缩模型的过程,突出了Claude在创意制作领域的应用,暗示了其在辅助视觉艺术和模型构建方面的潜力。
中文介绍 该视频探讨了人工智能模型实际「知道」什么的核心问题。内容可能涉及AI模型的知识边界、其学习和理解机制与人类认知的异同,以及AI系统如何获取、处理和表达信息。
中文介绍 这则由Claude发布的YouTube短视频,探讨了人工智能(AI)出现「幻觉」现象的原因。AI幻觉是指大型语言模型在生成文本时,提供看似合理但实际不准确或虚假信息的问题。该视频旨在解释为何AI会生成不符合事实的内容。
中文介绍 由Claude在YouTube发布的一则短视频,探讨了人工智能(AI)如何形成其“性格”。视频以此为主题,讨论AI系统在学习和训练过程中发展出独特行为模式的现象。
中文介绍 这段来自Claude(YouTube)的短视频,以“什么是阿谀奉承?”为题,旨在清晰阐释「阿谀奉承」(sycophancy)这一概念。视频内容可能涵盖该行为的定义、特征及其在人际互动中的表现形式,帮助观众深入理解这种社会现象的本质。
5 回复 · 程序员 节点
18 回复 · Apple 节点
15 回复 · Apple 节点
34 回复 · Apple 节点
17 回复 · Python 节点
5 回复 · Linux 节点
12 回复 · Apple 节点
43 回复 · Apple 节点
29 回复 · Apple 节点
6 回复 · Linux 节点
该源今日无内容。
44 points · 8 comments
151 points · 80 comments
41 points · 25 comments
84 points · 17 comments
32 points · 9 comments
16 points · 1 comments
30 points · 11 comments
55 points · 17 comments
66 points · 13 comments
131 points · 63 comments
112 points · 55 comments
245 points · 201 comments
84 points · 31 comments
93 points · 31 comments
284 points · 222 comments
153 points · 24 comments
241 points · 78 comments
645 points · 306 comments
83 points · 42 comments
329 points · 204 comments
363 points · 155 comments
126 points · 86 comments
140 points · 31 comments
15 points · 2 comments
25 points · 6 comments
5 points · 0 comments
15 points · 2 comments
125 points · 22 comments
99 points · 51 comments
74 points · 30 comments
What's changed Bug fixes and reliability improvements
中文介绍 Anthropic 旗下的 Claude Code 项目发布了 v2.1.220 版本。此次更新主要内容为错误修复和可靠性改进,旨在提升软件的稳定性和用户体验。
What's changed Added Claude Opus 5 (claude-opus-5), now the default Opus model — 1M context, fast mode at $10/$50 per Mtok Added sandbox.network.strictAllowlist setting to deny non-allowlisted hosts for sandboxed commands without prompting Added DirectoryAdded hook that fires after /add-dir or the S
中文介绍 Anthropic的Claude-code项目发布v2.1.219更新。此版本引入了Claude Opus 5(claude-opus-5)作为默认Opus模型,具备1M上下文。其快速模式定价为每百万token $10/$50。同时,更新新增“sandbox.network.strictAllowlist”设置,增强沙盒命令的网络安全。
What's changed Changed /code-review to run as a background subagent, so review work no longer fills your conversation and keeps stacked slash commands as its review target Added screen-reader announcements of deleted text for word and line deletions (Option+Delete, Ctrl+W, Cmd+Backspace, Ctrl+U, Ctr
What's changed Added emoji shortcode autocomplete in the prompt input: type :heart: to insert ❤️, or :hea for suggestions — disable with the emojiCompletionEnabled setting Added warnings when transcript writes are failing (e.g. disk full) or when session saving is off due to an inherited environment
What's changed Added sandbox.filesystem.disabled setting to skip filesystem isolation while keeping network egress control Fixed a slowdown in long sessions where message normalization cost grew quadratically with the number of turns, causing multi-second stalls and slow resumes Fixed auto mode deny
What's changed Claude no longer runs the /verify and /code-review skills on its own; invoke them with /verify or /code-review when you want them
What's changed Fixed single-segment dir/** allow rules like Edit(src/**) auto-approving writes to nested dir/ directories anywhere in the tree instead of only /dir Fixed a permission-check bypass affecting commands run in Windows PowerShell 5.1 sessions Fixed Bash permission checks to fail closed on
What's changed /fork now copies your conversation into a new background session (its own row in claude agents) while you keep working; the in-session subagent it used to launch is now /subtask Added claude auto-mode reset to restore the default auto-mode configuration, with a confirmation prompt (pa
What's changed Added --forward-subagent-text flag and CLAUDE_CODE_FORWARD_SUBAGENT_TEXT environment variable to include subagent text and thinking in stream-json output Fixed permission previews relayed to chat channels not neutralizing bidirectional-override, zero-width, and look-alike quote charac
What's changed Added a live elapsed-time counter to the collapsed tool summary line so long-running tool calls visibly tick instead of looking stuck Added a startup warning for Write(path), NotebookEdit(path), and Glob(path) permission rules — use Edit(path) or Read(path) instead Fixed isolation: 'w
Release 0.146.0-alpha.10.1
中文介绍 OpenAI Codex 发布了其项目的 rust-v0.146.0-alpha.10.1 版本更新。此次更新主要涉及软件版本迭代。
Release 0.146.0-alpha.11
中文介绍 OpenAI Codex 发布了其项目的 rust-v0.146.0-alpha.11 版本更新。此次更新主要涉及软件版本迭代。
Release 0.146.0-alpha.10
中文介绍 OpenAI Codex 发布了其项目的 rust-v0.146.0-alpha.10 版本更新。此次更新主要涉及软件版本迭代。
Release 0.146.0-alpha.9
中文介绍 OpenAI Codex 发布了其项目的 rust-v0.146.0-alpha.9 版本更新。此次更新主要涉及软件版本迭代。
Release 0.146.0-alpha.8
中文介绍 OpenAI Codex 发布了其项目的 rust-v0.146.0-alpha.8 版本更新。此次更新主要涉及软件版本迭代。
Release 0.146.0-alpha.7
中文介绍 OpenAI Codex项目发布了Rust语言版本的0.146.0-alpha.7更新。此为该项目的一个alpha阶段版本,通常包含开发中的新功能或修复。
Release 0.146.0-alpha.6
中文介绍 OpenAI Codex项目发布了Rust语言版本的0.146.0-alpha.6更新。此为该项目的一个alpha阶段版本,通常包含开发中的新功能或修复。
Release 0.146.0-alpha.3.1
中文介绍 OpenAI Codex项目发布了Rust语言版本的0.146.0-alpha.3.1更新。此为该项目的一个alpha阶段版本,通常包含开发中的新功能或修复。
Release 0.146.0-alpha.5
中文介绍 OpenAI旗下的Codex项目发布了软件的0.146.0-alpha.5版本更新。此次发布是一个alpha测试版,聚焦于Rust相关组件的迭代与改进,持续推进软件开发进展。
Release 0.146.0-alpha.4
中文介绍 OpenAI旗下的Codex项目发布了软件的0.146.0-alpha.4版本更新。此次发布是一个alpha测试版,聚焦于Rust相关组件的迭代与改进,持续推进软件开发进展。
今日AI领域聚焦于大模型性能的突破,Anthropic Opus 5 在智能基准测试中表现卓越;同时,AI安全与架构问题成为焦点,Hugging Face呼吁透明度应对自主代理攻击,而美国则在考虑对中国AI模型实施选择性禁令,凸显了AI技术快速发展下的伦理、安全与地缘政治挑战。
Anthropic 的 Claude Opus 5 在衡量真实智能的 ARC-AGI-3 基准测试中获得 30.2% 的分数,几乎是 GPT-5.6 Sol 此前 7.8% 记录的四倍,表现远超 Fable 5。开发者发现该模型能独立形成反射方程,展现了前所未见的智能行为。这一突破预示着大模型在理解与解决复杂通用智能任务方面的重大进展,为AI研究和应用设定了新高度。
`Kronos` 是一个专为金融市场语言深度定制的基础模型,旨在精准理解并处理金融新闻、财报、研报等复杂文本数据。它解决了传统通用模型在金融领域专业性不足的痛点,通过深入学习金融领域的独特术语和上下文,显著提升了市场情绪分析、信息提取和趋势预测的准确性。该模型将赋能金融分析师和量化研究员,帮助他们在高专业性领域做出更明智的决策。
由知名 AI 专家 Andrew Ng 推出的 aisuite 项目,旨在提供一个简洁统一的界面,用于访问和管理多个生成式 AI 供应商服务。它通过标准化 API 抽象层,解决了开发者在同时使用不同大语言模型或生成式 AI 平台时面临的接口不一致、集成复杂等问题。该平台使开发者能够更轻松地比较模型、快速切换供应商或构建多模态 AI 应用,从而提高开发效率。
Anthropic 的 Claude 大模型迎来重要更新,除了 Opus 5 版本正式推出外,其全新的语音功能「Claude Voice」被认为是更具影响力的进展。这一突破性技术将使 Claude 能够进行更自然的语音交互,极大地扩展了其应用场景,特别是在需要实时、口语化交流的领域。此举标志着大模型从文本到多模态交互的关键一步,为用户带来了全新的交互体验。
Aymo AI 是一个为团队设计的一体化人工智能平台,旨在整合各种AI工具和服务,提升团队协作效率和工作流程,赋能企业创新。该平台提供全面的AI解决方案,帮助团队实现智能化运营,通过集成多种AI功能,简化了团队管理和项目执行,致力于成为企业级AI应用的核心枢纽,推动工作效率和创新能力双提升。
在OpenAI遭遇「前所未有的」自主代理网络攻击后,Hugging Face的CEO呼吁采取「彻底的透明度」应对。他强调,这是首次由自主代理发起的网络攻击,需有同等规模的回应。此事件凸显了AI系统在安全领域面临的新挑战,以及行业内对于信息共享和协同防御的迫切需求,预示着AI安全将成为未来技术发展的重要考量。
据报道,特朗普政府出于安全考量,倾向于对中国AI模型实施有针对性的禁令,而非全面限制。尽管OpenAI和Google DeepMind公开反对开源模型监管,OpenAI和Anthropic却在私下为此游说。这一政策倾向反映出美国在AI领域的地缘政治考量,以及对中国AI技术发展的警惕,预示着未来AI国际合作与竞争的复杂性。
据Mark Gurman消息,苹果计划在明年6月的WWDC大会上揭示其首款智能眼镜,并预计在2027年底前推出。苹果将重点强调隐私保护,以区别于Meta引发争议的AI眼镜,作为其产品策略。这一举动表明苹果希望通过在AI硬件领域延续其对用户数据保护的承诺,以此在日益激烈的智能穿戴设备市场中建立竞争优势和用户信任。
本实用指南对当前 AI 智能体架构中常被混淆的「Agent Harness Engineering」、「Loop Engineering」和「Graph Engineering」三个概念进行了详细对比和区分。博主旨在厘清它们各自的特点及其对模型可靠性的影响,帮助读者深入理解不同架构层。文章为开发者和研究人员提供了构建高效、稳定 AI 智能体的思路,特别是在智能体设计模式不断演进的背景下。
这是 Anthropic 官方出品的 Claude 大型语言模型使用指南和示例代码集合。项目以 Jupyter Notebook 形式,展示了多种有趣且高效的 Claude 应用场景和 Prompt 工程技巧。旨在帮助开发者、研究人员或对大型语言模型感兴趣的用户,更好地理解和利用 Claude 的能力,加速基于 Claude 的应用开发和实验,从而提升模型应用的效率和创造性。
Cursor 的升级版代理群成功以 Rust 语言重建 SQLite,仅通过文档实现 100% 测试通过率,这表明了一种高效的 AI 编程新范式。该研究指出,当前沿大模型负责整体规划和策略制定时,成本较低的模型也能高效完成大部分具体的编程任务。这一发现为企业和开发者优化 AI 编程工作流、降低开发成本提供了新的思路,预示着未来 AI 辅助编程将更加普及和高效。
今天的 AI 产品发布聚焦于**Agent 应用的精细化落地**,尤其体现在开发工具、企业效率以及儿童教育等领域。多款创新产品揭示了 AI 智能体如何从概念走向各行业实际应用,以及 AI 在提升日常生产力方面的智能感知和个性化能力持续进化。
Yoggi 是一款专为儿童设计的安全人工智能聊天工具。它提供了一个受控且友好的 AI 对话环境,旨在帮助儿童以安全的方式探索 AI 互动,并进行学习和娱乐。该产品通过严格的内容过滤和隐私保护措施,解决了家长对 AI 工具潜在风险的担忧,为儿童提供了一个既能激发好奇心又能保障在线安全的智能伴侣,是当前市场稀缺的儿童友好型 AI 产品。
Aymo AI 是一个为团队设计的一体化人工智能平台,旨在整合各种 AI 工具和服务,显著提升团队协作效率和工作流程。该平台提供全面的 AI 解决方案,涵盖从内容生成到数据分析等多个方面,赋能企业创新。它解决了团队在使用分散 AI 工具时面临的集成复杂和效率低下问题,通过统一入口和协同功能,帮助各类团队实现智能化运营,是企业级 AI 部署的有效选择。
Openbase 是一款允许用户通过语音从任何地点管理其人工智能代理团队的工具。它旨在简化 AI 代理的协调与控制,提高团队在不同场景下的工作效率。该工具创新地结合了语音交互和远程管理能力,解决了传统 AI 代理管理界面复杂、操作不便的痛点。Openbase 特别适用于需要高度灵活和移动性工作的团队,使 AI 智能体能够真正融入日常工作流,实现无缝的智能团队管理。
citrolabs/ego-lite 是一款创新的浏览器,专为用户及其 AI 代理并行工作而设计。它解决了传统浏览器无法原生支持 AI 代理协同操作的痛点,使用户和 AI 智能体能够共同浏览网页、执行任务,显著提升工作效率。该项目零成本、零配置,并能共享登录状态而互不干扰,为 AI 代理进行研究、数据收集和自动化 Web 任务提供了强大且便捷的平台,让 AI 成为浏览器内的原生伙伴。
Chat2DB 是一款强大的 AI 驱动数据库工具和 SQL 客户端,提供直观的图形用户界面。它旨在简化数据库操作和管理,支持包括 MySQL、Oracle、PostgreSQL 在内的多种主流数据库。该工具通过 AI 能力辅助用户编写 SQL 查询、分析数据,有效解决了传统数据库客户端操作复杂、跨数据库兼容性差的问题。适用于开发者、DBA 和数据分析师,显著提升数据库交互和管理效率,让数据操作更智能。
alibaba/open-code-review 是阿里巴巴开源且免费的代码审查工具,经阿里大规模实践验证。它采用混合架构,结合确定性分析流水线和 LLM 智能体,能生成精确到行级别的代码评论。内置针对 NPE、多线程等问题的预训练规则集,旨在提升代码质量与开发效率。该工具解决了传统代码审查耗时耗力、依赖人工经验的痛点,为各类开发团队提供高效、智能化的代码审查解决方案。
aisuite 项目由知名 AI 专家 Andrew Ng 推出,旨在提供一个简洁统一的界面,用于访问和管理多个生成式 AI 供应商服务。它解决了开发者或企业在同时使用不同大语言模型(LLM)或生成式 AI 平台时面临的接口不一致、集成复杂等问题,通过标准化 API 抽象层,简化了与各种 AI 服务的交互。该项目使开发者能够更轻松地比较不同模型、快速切换供应商或构建多模态 AI 应用。适用于需要集成、测试或管理多个生成式 AI 模型的开发者、研究人员及企业。
BrainFeed 是一款个性化学习信息流工具,它能够智能地重定向用户的滚动行为,将注意力引导至更具价值的内容。该工具旨在通过提供相关且有益的知识,帮助用户更高效地获取信息,优化学习体验,解决信息过载和注意力分散的问题。它利用智能算法分析用户兴趣和行为,为其推荐高质量的学习资源,让每一次滚动都能带来有意义的收获。
PlayingFild 是一款理解上下文的生产力工具和标签页管理器。它旨在帮助用户更高效地组织和管理浏览器标签页,根据使用场景、工作流或特定项目智能地进行分类和切换,有效减少信息过载。该工具通过其上下文感知能力,解决了传统标签页管理工具缺乏智能、无法适应复杂工作模式的痛点,显著提高工作效率和专注度,为用户带来更流畅的浏览体验。
Shoplazza 推出的 Athena 是一款为整个商业堆栈提供服务的编排代理工具。它旨在自动化和优化电商运营中的各个环节,从库存管理、订单处理到营销策略,全面提升商家效率。通过部署智能 AI 代理,Athena 解决了电商运营复杂、人工成本高昂的痛点,帮助商家高效管理其商业生态系统,实现自动化决策和智能运营,从而专注于业务增长,适用于各类电商企业。
TouchGrass 是一款注重隐私的上下文感知休息提醒工具。它能够在不获取侵入性权限的前提下,根据用户的工作状态(如是否长时间集中于电脑屏幕)智能地提供休息提醒,帮助用户保持身心健康和工作效率。该工具解决了传统提醒应用机械、打扰的弊端,通过智能感知工作模式,以更人性化的方式鼓励用户适时休息,从而预防疲劳、提升整体幸福感,适用于长时间使用电脑的职场人士。