2026-09-03

Anthropic 发布电商 Agent 架构与生产实践指南,并开源 commerce-agents 参考实现 / 60 signals

Daily Edition

2026-09-0360 signals

Lead Story

3 items

Anthropic 发布电商 Agent 架构与生产实践指南,并开源 commerce-agents 参考实现

Anthropic 发布电商 Agent 构建指南,基于与零售、旅游、电信等团队的落地经验,核心架构是单个 Claude 在标准 Agent 循环中配合技能与工具,而非按领域拆分子智能体,并开源了 anthropics/commerce-agents 参考实现,含购物与商家 Agent。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtkcffah018zrog0zokk6s30

Read original ↗

Google AI 团队分享如何为 LLM-as-a-Judge 评测编写可靠的评分标准

Google AI 团队发布教程,讲解如何为 LLM-as-a-Judge 评测编写可靠的布尔式评分标准,指出模糊提示会导致评估不一致和浪费 token。文中给出四条经验:问题保持原子化且互不重叠、只让评判模型评估客观事实(可用 RFC 2119 术语如 MUST 表述)、只评 prompt 中明确要求的内容、用专家标注的 golden set 校准评判模型直至与人类评分一致。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtkbz92801nmrowy61g2fsob

Read original ↗

llm-gemini 0.34

Release: llm-gemini 0.34 New model gemini-3.8-flash for Gemini 3.8 Flash , with low, medium and high thinking levels. #146 Fixed async responses failing to record the resolved model version. Thanks, Charlie Tonneslan . #137 Google released Gemini 3.8 Flash (and 3.8 Flash Cyber, but that's available to "trusted defenders" only) today. Here are the pelicans for high, medium, and low. This is high: For comparison, here are the same pelicans generated using Gemini 3.7 Flash . Something I appreciate about Gemini Flash is that it's fast, cheap, and competent at things like HTML and JavaScript. I was messing around with it and prompted "make me a cool thing in html" and it built this , which is certainly a cool thing in HTML! Took 13 seconds, cost 1.8 cents. Your browser does not support HTML5 vi

Read original ↗

AI & Builders

8 items

Claude's new system prompt really doesn't want to reproduce song lyrics

Anthropic publish the system prompts for their Claude consumer applications ( Claude.ai and the Claude mobile apps - sadly not for Claude Cowork or Claude Code). I love that they do this, and that they share not just the current prompts but historic changes to their prompts as well. They used to keep all of the prompts on a single page, but when I checked today I noticed they had re-arranged those prompts into an index page and then a page per model - here's the page for Haiku 4.5 for example, which has the original prompt from October 15th 2025 and an updated prompt from January 18th 2026. A neat thing about Anthropic's platform.claude.com/docs site is that it's designed to be usable by LLMs. You can add .md to any page to get back the content as Markdown - here's the system prompt index

Read original ↗

Quoting Rick Brewster

Direct2D has always been the biggest hurdle for Paint.NET on WINE, and it's clear that it will never be completed enough for Paint.NET's use. And I can't just "disable" the use of Direct2D. So, instead, Paint.NET now has an internal, from-scratch, clean-room reverse-engineered rewrite of Direct2D that it uses on WINE (triggered by using /wine ). It lives in PaintDotNet.Windows.Direct2D1.Managed.dll . This was written by our good friend Claude , without whom this would NOT have been possible and would NEVER have happened. [...] Most of this code is, as they say, "vibe coded." By that I mean that it has not been thoroughly reviewed, it's more "trust me bro" style. I cannot possibly review 180,000 lines of code, it's just way way way too much. For reference, the rest of Paint.NET is about 700

Read original ↗

Quoting Andrew Digby

325 #kakapo! The chicks from this year's record breeding season are now juveniles and so have been added to the population. In 1995 there were just 51 kākāpō left. Recovery of critically endangered species is possible with sustained effort. — Andrew Digby , providing the best news of the year Tags: kakapo

Read original ↗

Tech & Industry

15 items

Researchers fear safety disaster ahead of OpenAI’s Astra release

OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it "may be the single worst development for AI security/safety to date." Shortly after […]

Read original ↗

Uber beats Waymo as first to launch robotaxis in London

Uber beat out Waymo in launching a commercial robotaxi service in London, the city's first. The vehicles use autonomous driving tech developed by Wayve, a UK-based startup, and will initially feature safety drivers behind the wheel. The launch is a milestone for Uber, which has been plotting a UK launch with Wayve for several years. […]

Read original ↗

NASA’s Cargo-Moving Robotic Arm Named 300th IEEE Milestone

In the 1960s NASA began developing a system of reusable space shuttles to make its work more efficient and to reduce costs. The shuttles could launch like rockets, maneuver in Earth’s orbit, and land like airplanes. They also could carry large satellites to and from orbit. Like other types of transportation, machinery eventually breaks down, and parts need to be replaced or fixed. And the cargo being carried to and from Earth has to be moved to its final destination. To complete such tasks, Spar Aerospace (now part of MDA Space ) of Brampton, Ont., Canada, and the National Research Council in Ottawa developed a robotic arm, the Shuttle Remote Manipulator System . The project was a joint venture between the U.S. and Canadian governments. Known as Canadarms , the robotic tools attached to sh

Read original ↗

Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake’

The internet has a trust problem, and it’s not just because social media feeds are filling up with AI slop. AI-generated text and images are now making their way into job applications, product reviews, and even insurance claims, leaving platforms and users alike scrambling to figure out what’s real. A handful of startups have cropped up in the past couple of […]

Read original ↗

The best tech and gadgets announced at IFA so far

The doors to Europe's largest consumer tech show haven't opened to the public yet, but there's already plenty of news coming out of IFA 2026 in Berlin, Germany. If you're already struggling to keep up with what has been announced, here are some of the best new gadgets and upgrades from the show - including […]

Read original ↗

1Password wades into a right-wing mess after funding a Linux project

1Password faced immediate backlash from customers this week over a $300,000 pledge in support of a Linux distro created by David Heinemeier Hansson, who has regularly published overtly racist blog posts that include comments calling for deportation of ethnic minorities in Europe. The popular password manager is now a "distinguished corporate patron" of Omacom, the […]

Read original ↗

Amazon’s AI assistant can now spot fake emails from the company

Amazon is trying to combat impersonation scams with a new feature that allows you to use its AI assistant to determine whether an email, text message, or phone call actually came from the company. With the update, you can ask Alexa for Shopping about a message you received, and it will use AI to compare […]

Read original ↗

AI Travel & Travel Tech

10 items

World Affairs

15 items

Design & Culture

1 items

Outside the Bubble

1 items

Builders on X

7 items

AI for cyber is about to go vertical. The models increasingly becoming insanely good at finding and exploiting vulnerabilities. Frontier models are ahead, but we’re already seeing that open weights is not far behind. Most enterprises are already inundated with cyber discoveries, so now that’s only going to multiply. Triaging and automating the fixes with more AI -along with human oversight- is essentially the only way forward. In case you were wondering what jobs AI was going to create, it’s definitely going to great time to be in security.

Open on X ↗

Terrifying.. also just the start of how leaks like this are going to get way worse 😭 https://t.co/gN8PiNCUXJ

Open on X ↗

I wish every city was in Tokyo https://t.co/j1ppDGCx7p

Open on X ↗

fwiw I dislike having alot of random skills installed. I'm at a point where I only have about a dozen or so (mostly my own) and I regularly delete skills that I don't use anymore. Also I try to keep all my skills as short as possible. I encourage you to do the same.

Open on X ↗

While I'm cleaning up my AI skills I have a question for the experts here. Sometimes I end up in this pattern: 1. I run my skill and it doesn't get it perfect in one shot 2. I do manual iteration with AI to get it right 3. I then ask AI something like: "How would you update the skill so we can one-shot this?" before reviewing and approving its edits. The problem is AI often overfits on this one thread and over time the skill drifts. Any good solutions here?

Open on X ↗