AI News Aug 20, 2026
By Frontier Editorial •
Key Takeaways
- OpenAI lays out new security changes after its AI hacked Hugging Face: OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environmen…
- OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous: OpenAI is deliberately "pacing AI model development," partly because the upcoming "Astra" …
- OpenAI institutes new safeguards after Hugging Face breach: The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasi…
- Hacker News: Reported (community)
- Enterprise teams running AI agents at scale are finding that a single model handles every task poorly — either the model is too expensive for simple questions o…
What are the top AI breakthroughs?
This Aug 20, 2026 covers 180 curated AI news items spanning technology, research, and product developments. TrueFoundry's open source AI agent harness TrueForge boasts 30%-75% cheaper task completion than Claude Managed Agents: ...
TrueFoundry's open source AI agent harness TrueForge boasts 30%-75% cheaper task completion than Cla…
TrueFoundry's open source AI agent harness TrueForge boasts 30%-75% cheaper task completion than Claude Managed Agents: Another day, another new AI agent harness is released. Only this time, it's one that aims to solve a growing enterprise problem as AI agents proliferate: enabling greater developer control of agents and tools, while reducing cost. TrueFoundry , a San Francisco B2B machine learning startup co-founded in 2021 by former Meta engineers, has released its own c...
Mojo🔥 is now open source: Mojo🔥 is now open source Mojo🔥 is now open source The Mojo programming …
Mojo🔥 is now open source: Mojo🔥 is now open source Mojo🔥 is now open source The Mojo programming language has been promising an open source release since May 2023 . Last week they shipped their 1.0 and today they have followed through on that original promise, releasing the compiler and toolchain under an Apache 2 license. When Mojo first launched the stated goal was to produce...
Offering Zero Data Retention for frontier models: OpenAI reaffirms Zero Data Retention for eligible …
Offering Zero Data Retention for frontier models: OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
Offering Zero Data Retention for frontier models: OpenAI reaffirms Zero Data Retention for eligible …
Offering Zero Data Retention for frontier models: OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
OpenSourcing TrueForge Agent harness : Expecting feedback from community on the agent loop: Hey folk…
OpenSourcing TrueForge Agent harness : Expecting feedback from community on the agent loop: Hey folks 👋 We just open sourced TrueForge, our vendor-neutral agent harness for building general-purpose agents. It handles the runtime pieces that get painful quickly : context management, tool/MCP execution, subagents, sandboxing, approvals, persistent state, and more. We also benchmarked the harness itself. With the same Opus 4.8 model, TrueForge del...
GLM-5.3 tops the open-model rankings and undercuts rivals on price, but its release is delayed: GLM-…
GLM-5.3 tops the open-model rankings and undercuts rivals on price, but its release is delayed: GLM-5.3, the AI model from Chinese startup Z.ai, scores 60 points on the Artificial Analysis Intelligence Index. That ties it with Kimi K3 for the top spot among open models, and it's seven points ahead of the previous GLM-5.2. The article GLM-5.3 tops the open-model rankings and undercuts rivals on price, but its release is delayed appeared first on The...
GLM-5.3 hits the API at $1.4/$4.4 per million tokens: After a stunning debut last week with cyber ca…
GLM-5.3 hits the API at $1.4/$4.4 per million tokens: After a stunning debut last week with cyber capabilities so advanced they reportedly found a previously undetected vulnerability in Cursor, GLM-5.3, the new frontier open source language model from Chinese startup z.ai, has now hit the application programming interface (API) — allowing developers the ability to build atop it and plug it into their agents...
Block’s new Apache 2.0 agent workspace Berd works across models and harnesses, stores conversation h…
Block’s new Apache 2.0 agent workspace Berd works across models and harnesses, stores conversation history locally: Block , the technology company founded by former Twitter CEO Jack Dorsey that owns Square, Cash App and the music streaming service Tidal, is open-sourcing Berd , a desktop application it originally built to give its own employees a single environment for working with AI agents across different models, tools and projects. Berd is a locally installed graph...
Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required: The bigg…
Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required: The biggest AI model release of the past few days, at least among the developers and AI power users on social media, wasn't a frontier cloud model from OpenAI, Anthropic or Google. It was a 27-billion-parameter model from Alibaba: Qwen3.8-27B landed on Hugging Face on Friday under an enterprise-friendly, open source Apache 2.0 license, giving developers d...
smolmachines / smolvm as a sandbox for untrusted Python & JavaScript: Research: smolmachines / smolv…
smolmachines / smolvm as a sandbox for untrusted Python & JavaScript: Research: smolmachines / smolvm as a sandbox for untrusted Python & JavaScript I tasked Claude Fable 5 running in Claude Code for web with the following research task: Put https://smolmachines.com through its paces as a fast secure sandbox. Explore what it would take to use this to run untrusted Python and JavaScript code in a way that is limited in what...
Anthropic says any lab can now let a language model agent run the whole protein design stack: Anthro…
Anthropic says any lab can now let a language model agent run the whole protein design stack: Anthropic had its Claude models design small proteins on their own that dock onto target structures in the body, a key step in early drug development. The hit rate reached up to 35 percent, far above the industry average of 10 to 15 percent. Claude only steered existing specialized tools, and an independent review is still pending. The article Anthropic s...
Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race: Cur…
Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race: Cursor began rolling out Origin , its own code hosting platform, to paid users on Monday morning. Roughly three and a half hours later, GitHub's status page lit up with what became a six-hour-and-forty-two-minute global degradation — error rates near 20% across pull requests, issues and the API, and near 50% on archive and raw file downloads, according to...
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things: Friday's big release was Q…
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things: Friday's big release was Qwen 3.8 27B , an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor Qwen 3.6 27B was impressive. Qwen's self-reported benchmarks for this model are eye-opening....
Strengthening democratic oversight in national security: OpenAI launches an initiative to strengthen…
Strengthening democratic oversight in national security: OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools, training, and expertise.
Replit expands access to software creation with GPT-5.6 Luna: Replit introduces Free Mode, powered b…
Replit expands access to software creation with GPT-5.6 Luna: Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.
OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous: OpenAI is …
OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous: OpenAI is deliberately "pacing AI model development," partly because the upcoming "Astra" model may be close to gaining critical cyberattack capabilities. A new monitoring system triggers an alert within 30 minutes if a model shows suspicious behavior. The article OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous app...
New benchmark ranks search APIs for AI agents on quality, cost, and speed: Artificial Analysis has r…
New benchmark ranks search APIs for AI agents on quality, cost, and speed: Artificial Analysis has released the "Search Index," a benchmark that rates search API providers for AI agents on quality, cost, and speed. Of seven providers tested with GPT-5.6 Luna, Parallel, Exa, and Firecrawl scored highest. The article New benchmark ranks search APIs for AI agents on quality, cost, and speed appeared first on The Decoder .
Pacing model development in an era of cyber-critical capabilities: OpenAI is strengthening monitorin…
Pacing model development in an era of cyber-critical capabilities: OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
OpenAI lays out new security changes after its AI hacked Hugging Face: OpenAI is announcing security…
OpenAI lays out new security changes after its AI hacked Hugging Face: OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have "critical" cybersecurity capa...
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one: …
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one: Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows . In July, 13% of 108 enterprises surveyed said they trust automated evaluation, u...
LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
ChatGPT Ads expands across Europe: ChatGPT Ads is expanding to 31 European markets. Learn how advert…
ChatGPT Ads expands across Europe: ChatGPT Ads is expanding to 31 European markets. Learn how advertisers can reach people as they explore, compare options, and make decisions.
Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn: The NSA…
Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn: The NSA, CISA, and FBI say attackers are using AI to build exploit scripts targeting Siemens S7 controllers, drastically cutting the time and skill needed to attack industrial control systems. Critical U.S. sectors like energy, water, and manufacturing are affected. The article Attackers are using AI to build exploits for industrial control systems, U.S....
Unsloth Dynamic 3.0 GGUFs
Unsloth Dynamic 3.0 GGUFs
Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed E…
Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape model behavior, but it does not create an execution boundary. We introduce Aegis, a runtime governance system...
ASI-Bench: At the Dawn of Artificial Superintelligence: Artificial superintelligence (ASI) requires …
ASI-Bench: At the Dawn of Artificial Superintelligence: Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchm...
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index: Qwen 3.8 27B scores 52 on the …
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index: Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.7T parameters , and Luna is size unknown but presumably a whole lot bigger than 27B. Qwen 3.8 27B is a truly astonishing model . V...
OpenAI fixes Codex bug that deleted real user files without permission: OpenAI patched Codex after G…
OpenAI fixes Codex bug that deleted real user files without permission: OpenAI patched Codex after GPT-5.6 Sol started deleting real user files on its own. A cleanup command meant for temporary folders was wiping home directories instead. Codex now verifies deletion targets first, and full-access mode can no longer be triggered by accident. The article OpenAI fixes Codex bug that deleted real user files without permission app...
Meta AI is getting a Mac app: Meta is launching a new Mac app dedicated to its AI chatbot. In an ann…
Meta AI is getting a Mac app: Meta is launching a new Mac app dedicated to its AI chatbot. In an announcement on Wednesday, Meta says you can share your window with its AI chatbot, which can provide suggestions, answer questions, or create content based on what's on your screen. Meta AI on the Mac also supports dictation across all apps. The […]
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models: Safety alignm…
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models: Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff...
I built a visual architecture & token-reduction diagram engine for multi-agent LLM pipelines: When w…
I built a visual architecture & token-reduction diagram engine for multi-agent LLM pipelines: When working with multi-agent LLM systems, the hard part usually isn't getting a response—it's knowing what actually happened under the hood: which model handled what, what was sent over the network, how much it cost, and whether sensitive data was masked before leaving your machine. To solve this, I added a visual diagram engine to **Mova Context** in th...
Anthropic CEO says AI centralizes by nature and open models just shift power to whoever owns the chi…
Anthropic CEO says AI centralizes by nature and open models just shift power to whoever owns the chips: An open fight over AI regulation has broken out on X. Investor Gavin Baker, former White House adviser David Sacks, and Meta researcher Yann LeCun accuse Anthropic CEO Dario Amodei of using fear rhetoric to buy himself a regulatory advantage. Amodei counters that regulation can also rein in corporate power, and that open models alone just shift power towa...
The Defender’s Window: AI is reshaping cybersecurity for attackers and defenders alike. Learn how Op…
The Defender’s Window: AI is reshaping cybersecurity for attackers and defenders alike. Learn how OpenAI is strengthening its defenses and what security teams can do now.
DFlash 2: Keep Drafting Parallel
DFlash 2: Keep Drafting Parallel
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models: Small…
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these b...
Partnering with CodeAI to prepare the first AI generation: OpenAI and CodeAI are partnering to help …
Partnering with CodeAI to prepare the first AI generation: OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it responsibly.
A Local Opus? Alibaba Qwen Open-Sources Qwen3.8-27B — Frontier Coding and Agent Scores That Runs on …
A Local Opus? Alibaba Qwen Open-Sources Qwen3.8-27B — Frontier Coding and Agent Scores That Runs on 17GB of RAM: Alibaba's Qwen team open-sourced Qwen3.8-27B, which topped Hugging Face's global trending chart within two days and passed one million downloads. Overseas developers nicknamed it the local Opus 4.6: under 30B parameters, it outperforms every model released four months ago including Opus 4.6, matches DeepSeek V4-Pro and GPT 5.6 Luna, and runs on 17GB of RA...
Alibaba Cloud's Ambition Is Not Agent Builder: Agent Studio Becomes an All-in-One Enterprise Agent S…
Alibaba Cloud's Ambition Is Not Agent Builder: Agent Studio Becomes an All-in-One Enterprise Agent Stack: Alibaba Cloud upgraded its agent services into Agent Studio, an all-in-one enterprise agent full-stack platform launched on Bailian at the August 14 Apsara release. The platform targets the dirty work of agent infrastructure: managed runtime, unified API keys across MCP services, agentic search, and memory, as cloud vendors race to own the agent runtime l...
Researchers say OpenAI revoked their access to limited cyber program: The idea behind OpenAI's Trust…
Researchers say OpenAI revoked their access to limited cyber program: The idea behind OpenAI's Trusted Access for Cyber program is to give trusted defenders better models so they can report bugs and vulnerabilities to companies, with the aim of getting flaws patched faster.
OpenAI hit the brakes. Now what?: With a looming IPO, intense competition from Anthropic, and Chines…
OpenAI hit the brakes. Now what?: With a looming IPO, intense competition from Anthropic, and Chinese and open-weight rivals nipping at its heels, OpenAI has plenty of reasons to move fast. Instead, it hit the brakes. On Tuesday, the company said it had slowed the pace of some AI development while it tightened security and safeguards. That included a two-week pause […]
Ornith-1.5: From Self-Scaffolding to Self-Improvement
Ornith-1.5: From Self-Scaffolding to Self-Improvement
China lets Nvidia's H200 chips trickle onto the mainland to help its AI firms keep pace with the US:…
China lets Nvidia's H200 chips trickle onto the mainland to help its AI firms keep pace with the US: China is letting small batches of Nvidia's H200 chips onto the mainland to help domestic AI firms in the race with the US. The article China lets Nvidia's H200 chips trickle onto the mainland to help its AI firms keep pace with the US appeared first on The Decoder .
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration: Long-context prefill…
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration: Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial precision routing over hardware-aligned score tiles outside fused dense attention. We...
Introducing ChatGPT for Teens: Built for learning, backed by protections: ChatGPT for Teens helps te…
Introducing ChatGPT for Teens: Built for learning, backed by protections: ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features, and additional controls for parents.
人人都能写好听的歌,阿里发布AI音乐模型HappyShrimp: 8月17日,阿里巴巴发布AI音乐模型HappyShrimp
人人都能写好听的歌,阿里发布AI音乐模型HappyShrimp: 8月17日,阿里巴巴发布AI音乐模型HappyShrimp
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge: DeepSeek's V4 Flash…
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge: DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step t...
Google Gemini is getting a dedicated student hub: As we're gearing up for back-to-school season, Goo…
Google Gemini is getting a dedicated student hub: As we're gearing up for back-to-school season, Google is rolling out a new dedicated student hub in Gemini. It's a one-stop repository for collecting research in a study notebook, creating flashcards, taking practice quizzes, and more. Google is also enhancing its study notebooks with support for graphs and images. It can even add test dates […]
OpenAI is testing Private Safety Processing, a new technique to identify misuse patterns while prese…
OpenAI is testing Private Safety Processing, a new technique to identify misuse patterns while preserving zero data retention protections, with early customers (Ina Fried/Axios): Ina Fried / Axios : OpenAI is testing Private Safety Processing, a new technique to identify misuse patterns while preserving zero data retention protections, with early customers — OpenAI said Wednesday that it believes a new technique will allow it to safely serve its most advanced models to businesses without needing to retain their data.
Unsloth Desktop Review: The Local AI App I’d Use Beyond Chat: Unsloth Desktop brings local models, a…
Unsloth Desktop Review: The Local AI App I’d Use Beyond Chat: Unsloth Desktop brings local models, agent tools, synthetic-data workflows, fine-tuning, and exports into one polished interface. Here’s why it stands out.
全球首个人形机器人自主乒乓球完整对局亮相2026世界机器人大会: 超维动力KAI全栈具身智能硬核登场
全球首个人形机器人自主乒乓球完整对局亮相2026世界机器人大会: 超维动力KAI全栈具身智能硬核登场
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LL…
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking: Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similarity, which can favor familiar ideas over novel ones. We propose a logit-based energy s...
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs: Group-relative policy …
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories differ substantially in their interaction efficiency. For instance, circuitous successe...
AI Rebuilds Consumer Markets: ByteDance Large Models Unlock New Consumption Increment: ByteDance's a…
AI Rebuilds Consumer Markets: ByteDance Large Models Unlock New Consumption Increment: ByteDance's answer is embedding large-model capabilities as infrastructure into enterprise workflows: Doubao's scheduled agents generate competitor briefings, Seedance 2.5 synthesizes physical-world training data, and Doubao 2.1 Pro handles production-grade coding. Daily token calls passed 180 trillion in June 2026, up over 10x year over year.
Everything That Happened in AI Today (Tuesday, August 18, 2026): OpenAI kept its largest planned fro…
Everything That Happened in AI Today (Tuesday, August 18, 2026): OpenAI kept its largest planned frontier RL run on hold as cyber safeguards tightened; Google won Spirit Airlines’ data auction; Etched hit a $21B valuation; physical-AI funding reached $47.4B; Axiom formally verified the BGP246 prime-gap theorem.
How NVIDIA scales expertise with ChatGPT Work: NVIDIA teams use ChatGPT Work to reduce manual tasks,…
How NVIDIA scales expertise with ChatGPT Work: NVIDIA teams use ChatGPT Work to reduce manual tasks, connect fast-moving signals, and scale successful workflows globally.
GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerabi…
GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerability' in Cursor: Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, today released GLM-5.3 with substantial gains in long-horizon coding and a more consequential — and potentially sensitive — jump in cybersecurity capabilities. Already, GLM-5.3's cyber capabilities have found a "potentially...
Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't t…
Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done: Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other's Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival's work. There was no prompt injection an...
Meta launches a Mac app for Meta AI and says Meta AI can now work directly with Instagram and Facebo…
Meta launches a Mac app for Meta AI and says Meta AI can now work directly with Instagram and Facebook accounts, Meta ad campaigns, and Google Workspace (Emma Roth/The Verge): Emma Roth / The Verge : Meta launches a Mac app for Meta AI and says Meta AI can now work directly with Instagram and Facebook accounts, Meta ad campaigns, and Google Workspace — You can share your window with the new Meta AI app, as well as connect it to Google Workspace. … Meta is launching a new Mac app dedicated to its AI chatbot.
Guidelight’s First AI Control Grades Put Anthropic and OpenAI on Top: The new nonprofit Guidelight g…
Guidelight’s First AI Control Grades Put Anthropic and OpenAI on Top: The new nonprofit Guidelight graded Anthropic, OpenAI, Google, xAI, and Meta on six agent-control practices. Its first scorecard finds meaningful progress in monitoring, but less evidence that labs can consistently block or contain unsafe actions.
From smart cockpits to AI-native cars, Banma Intelligence eyes the next wave of automotive software:…
From smart cockpits to AI-native cars, Banma Intelligence eyes the next wave of automotive software: As large AI models accelerate their integration into vehicles, the competitive dynamics of intelligent cars are changing. In the past, smart cockpits largely focused on voice assistants, in-car applications, and multimedia services. Today, with on-device omni-models, AI agents, and AI operating systems gradually becoming reality, cars are evolving from sm...
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents: Clinical tr…
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents: Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulatory submissions, yet LLM-based code generation fails catastrophically on this task: across 11 single-shot attempts with five frontier models, none produces a valid subject-level analysis dataset. We introduce GxP-Agent...
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimizatio…
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or i...
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents: Reinforcement learning for coding …
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents: Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while...
OpenAI says the changes to its model training will increase compute overhead by 20% of observed infe…
OpenAI says the changes to its model training will increase compute overhead by 20% of observed inference workload; the increase will not be handed to customers (Thomas Claburn/The Register): Thomas Claburn / The Register : OpenAI says the changes to its model training will increase compute overhead by 20% of observed inference workload; the increase will not be handed to customers — Expanded multistage chain of thought monitoring makes frontier model work more expensive — OpenAI on Tuesday said its decision …
Cursor capitalizes on GitHub frustration, launches rival hosting platform: Cursor, known for its AI …
Cursor capitalizes on GitHub frustration, launches rival hosting platform: Cursor, known for its AI Code Editor, is launching a new code-hosting platform to rival developers' long preferred favorite, GitHub.
OpenAI launches a ChatGPT version built for teens: OpenAI is shipping a version of ChatGPT tailored …
OpenAI launches a ChatGPT version built for teens: OpenAI is shipping a version of ChatGPT tailored to users aged 13 to 17. The article OpenAI launches a ChatGPT version built for teens appeared first on The Decoder .
Z.ai launches GLM-5.3 with claimed 50% gain on coding benchmark: Z.ai, the international brand of Ch…
Z.ai launches GLM-5.3 with claimed 50% gain on coding benchmark: Z.ai, the international brand of Chinese AI company Zhipu, has launched GLM-5.3, an update focused on coding, long-horizon tasks and cybersecurity. The model uses the same base model as GLM-5.2, with the company attributing the latest gains to post-training. Z.ai said GLM-5.3 scored 50% higher than GLM-5.2 on its internal Z.ai Code Bench and reached […]
Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut: Google is rol…
Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut: Google is rolling out Gemini 3.7 Flash , a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API prices in half. The release arrives just three weeks after the release of Gemini 3.6 Flash , an unusually short turnaround that Google attributes to developer f...
OpenAI seeks to one-up Anthropic with new customer privacy protections: A competition is developing …
OpenAI seeks to one-up Anthropic with new customer privacy protections: A competition is developing between OpenAI and Anthropic over who can provide the best privacy protections for enterprise customer data.
Anthropic’s Project Parka sits through meetings and assigns Claude agents the homework: submitted by…
Anthropic’s Project Parka sits through meetings and assigns Claude agents the homework: submitted by /u/ryanmerket [link] [comments]
IDC发布2026中国AI50强:360以“智能体+安全”双轮驱动入选: 凭借企业级智能体与AI安全的全栈布局,360成为中国人工智能产业发展的代表企业之一。
IDC发布2026中国AI50强:360以“智能体+安全”双轮驱动入选: 凭借企业级智能体与AI安全的全栈布局,360成为中国人工智能产业发展的代表企业之一。
I brought an ancient Zen book to life with Claude: I just released an interactive edition of The Gat…
I brought an ancient Zen book to life with Claude: I just released an interactive edition of The Gateless Gate, a collection of 49 Zen koans from 1228. Every koan gets its own 3D scene, with its own soundscape and a full spoken reading. Other than the speech, it's all generated procedurally on startup, so there's nothing to download. Demo: https://killedbyapixel.github.io/GatelessGate/ It started as a tes...
π0引用的中国团队,又出手了:世界仿真器新作发布: 给机器人造一个更接近真实的“第二世界”
π0引用的中国团队,又出手了:世界仿真器新作发布: 给机器人造一个更接近真实的“第二世界”
DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher p…
DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices: DeepSeek is expanding beyond the model layer and deeper into the software developers use to put AI agents to work. The Chinese AI lab on Thursday launched the official version of DeepSeek-V4-Pro , an updated flagship model focused heavily on agentic workloads, alongside DeepSeek Harness v0.1 , a new open-source agent harness that gives developers an alter...
Meet the startup helping Wall Street put a price on AI compute: The AI buildout shows no signs of sl…
Meet the startup helping Wall Street put a price on AI compute: The AI buildout shows no signs of slowing. And with hundreds of billions of dollars a year going into data centers and GPUs, compute has become the single biggest cost for anyone building AI products. But for all that spending, there still isn’t a straightforward way to put a price on compute — or for firms to hedge their exposure when the price changes....
Memo: the GOP asks AI companies to stem the public backlash against data centers, saying the issue c…
Memo: the GOP asks AI companies to stem the public backlash against data centers, saying the issue could kill its chances of holding a vital Senate seat in Ohio (Alex Isenstadt/Axios): Alex Isenstadt / Axios : Memo: the GOP asks AI companies to stem the public backlash against data centers, saying the issue could kill its chances of holding a vital Senate seat in Ohio — The Senate GOP campaign arm, in a private memo to top AI companies, warns that toxic views of U.S. data centers are killing …
章鱼动力亮相WRC 2026,携“脑-手-数据”技术体系探索具身智能未来范式
章鱼动力亮相WRC 2026,携“脑-手-数据”技术体系探索具身智能未来范式
WRC 2026 Opens: 34 Guangdong Companies March In, Showcasing a Complete Robot Supply Chain: The 11th …
WRC 2026 Opens: 34 Guangdong Companies March In, Showcasing a Complete Robot Supply Chain: The 11th World Robot Conference opened August 19 at Beijing E-Town with over 300 exhibitors, 150-plus new releases, and a first-ever procurement day. Guangdong sent 34 companies spanning the full robot supply chain, from EVE Energy batteries and ORBBEC 3D vision to UBTECH and LimX Dynamics full machines, as 49 central state-owned enterprises joined for th...
具身数据底座开卖,首发5100元:机器人训练数据有了新解法: 构建全栈物理AI基础设施
具身数据底座开卖,首发5100元:机器人训练数据有了新解法: 构建全栈物理AI基础设施
Alibaba Cloud adds a third data center in South Korea: Alibaba Cloud has brought its third data cent…
Alibaba Cloud adds a third data center in South Korea: Alibaba Cloud has brought its third data center in South Korea online, adding enterprise cloud services and six Agentic AI services for local customers, according to the announcement. The new site expands Alibaba Cloud’s infrastructure in one of Asia’s major technology markets. The company said its global network will reach 104 availability zones, adding...
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations: LLM ag…
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations: LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource w...
Anthropic details two experiments showing how Claude can accelerate protein design and analytical ch…
Anthropic details two experiments showing how Claude can accelerate protein design and analytical chemistry, and says it plans an access program for scientists (Anthropic): Anthropic : Anthropic details two experiments showing how Claude can accelerate protein design and analytical chemistry, and says it plans an access program for scientists — Summary: In this post, we share two results that show how Claude can help life scientists increase the pace of their research.
Sam Altman says OpenAI's decision to pace its AI development was caused by a collection of research …
Sam Altman says OpenAI's decision to pace its AI development was caused by a collection of research observations showing "various degrees of misalignment" (Alex Heath/Time): Alex Heath / Time : Sam Altman says OpenAI's decision to pace its AI development was caused by a collection of research observations showing “various degrees of misalignment” — “I think it is a good time to slow down,” OpenAI CEO Sam Altman told me last week, describing the company's decision …
OpenAI institutes new safeguards after Hugging Face breach: The new safeguards include more detailed…
OpenAI institutes new safeguards after Hugging Face breach: The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process.
AI systems quietly drop user instructions when they compress context: When AI systems condense long …
AI systems quietly drop user instructions when they compress context: When AI systems condense long conversations, they drop an average of 83 percent of user rules, like "don't send emails without my approval." Penn State researchers propose a small add-on module built on Qwen3.5-9B that preserves over 90 percent of these restrictions. The article AI systems quietly drop user instructions when they compress context appeared...
DeepSeek Harness Open Source: Everything Is a Plugin — the Bet Is an Agent Platform, Not a Product: …
DeepSeek Harness Open Source: Everything Is a Plugin — the Bet Is an Agent Platform, Not a Product: DeepSeek open-sourced DeepSeek Harness (CLI: dsh) on August 13, and GitHub stars passed 140,000 within days. Built on the Cordis microkernel, every component including the agent loop itself is a plugin, betting that Harness becomes the baseboard of the agent era: Model + Harness = Agent.
Claude Sonnet 5 shifts behavior when it recognizes the user as an AI safety researcher: submitted by…
Claude Sonnet 5 shifts behavior when it recognizes the user as an AI safety researcher: submitted by /u/rhiever [link] [comments]
I built an MCP server that lets two Claude Code sessions on different machines message each other: i…
I built an MCP server that lets two Claude Code sessions on different machines message each other: i kept copy pasting between two machines regarding APIs and architecture. between my PC's Claude Code which was supposed to work on my frontend and my other claude code sessions running on my ubuntu VPS was working on the backend. So I built an intercom which used channels API of anthropic as well as MCP server to transmit messages between 2 claude code s...
MiniMax核心工程负责人阿岛离职: 从技术研发到开发者沟通,长期活跃在开发者一线
MiniMax核心工程负责人阿岛离职: 从技术研发到开发者沟通,长期活跃在开发者一线
LLM-Only PDDL Domain Repair with Open-Weight Models: AI planning is concerned with finding a sequenc…
LLM-Only PDDL Domain Repair with Open-Weight Models: AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such models can be detected and repaired. For example, users may provide positive test plans tha...
Fresh ChatGPT chats are fast. Established ones now take 10–40 minutes or fail (HAR data): Long post,…
Fresh ChatGPT chats are fast. Established ones now take 10–40 minutes or fail (HAR data): Long post, but there’s a TL;DR first. I’m writing it this way because a vague “ChatGPT is slow” post wouldn’t be useful. I’ve also included the HAR measurements for anyone who wants the technical details. TL;DR Since around August 17–18 , my established ChatGPT conversations have suddenly become dramatically slower and much less reliable. I’m not only tal...
网易传媒发布”蜜蜂AI” :从工具到伙伴,让AI更懂人: 8月18日,网易传媒举办“蜜蜂AI媒体沟通会”
网易传媒发布”蜜蜂AI” :从工具到伙伴,让AI更懂人: 8月18日,网易传媒举办“蜜蜂AI媒体沟通会”
Alibaba launches HappyShrimp 1.0 AI music model: Alibaba has officially launched HappyShrimp 1.0, an…
Alibaba launches HappyShrimp 1.0 AI music model: Alibaba has officially launched HappyShrimp 1.0, an AI music model that can turn emotions, stories or memories into complete music tracks through natural-language prompts. The launch moves HappyShrimp from an earlier reported project into a publicly available product. Alibaba has not disclosed detailed information about the model’s technical architecture,...
AI Used to Verify Toughest Mathematics Proof Yet: Representing a significant milestone in AI-assiste…
AI Used to Verify Toughest Mathematics Proof Yet: Representing a significant milestone in AI-assisted mathematical research, a team at Axiom Math has automatically verified the proof of a theorem relating to prime numbers—colloquially referred to as the “246 theorem”—for the first time using the company’s AI system AxiomProver. In formal verification, mathematicians task a computer with checking a machin...
Context Engineering vs Prompt Engineering: www.elastic.co/search-labs/blog/context-engineering-vs-pr…
Context Engineering vs Prompt Engineering: www.elastic.co/search-labs/blog/context-engineering-vs-prompt-engineering submitted by /u/AvenueJay [link] [comments]
AI for Science开始“动手”了:机器人正式走进国家级实验室: 源络科技,推动AI for Science走进实验室3.0
AI for Science开始“动手”了:机器人正式走进国家级实验室: 源络科技,推动AI for Science走进实验室3.0
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract: API buyers purchase a date…
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort...
Cerebras unveils CS-4, a server rack powered by three WSE-3 Turbo chips and built around its new Nex…
Cerebras unveils CS-4, a server rack powered by three WSE-3 Turbo chips and built around its new Nexus architecture, with first shipments starting this quarter (Max A. Cherney/Reuters): Max A. Cherney / Reuters : Cerebras unveils CS-4, a server rack powered by three WSE-3 Turbo chips and built around its new Nexus architecture, with first shipments starting this quarter — Cerebras Systems (CBRS.O) announced on Tuesday a new version of its server hardware that includes its dinner-plate-sized chips that it says will speed AI chatbot queries.
As AI beats doctors, regulators shouldn't force a human into the loop, JAMA piece says: An opinion p…
As AI beats doctors, regulators shouldn't force a human into the loop, JAMA piece says: An opinion piece in the medical journal JAMA argues that autonomous AI will soon outperform any doctor-AI team at medical reasoning tasks. The authors warn against writing a doctor's final say into regulation, but concede that almost all the evidence comes from simulations, not real patient care. The article As AI beats doctors, regulators shouldn't force...
ChatGPT is getting a dedicated mode for teens: OpenAI is introducing a dedicated ChatGPT mode for te…
ChatGPT is getting a dedicated mode for teens: OpenAI is introducing a dedicated ChatGPT mode for teenagers, combining existing youth safeguards and new safety features under one roof. The launch comes amid mounting public scrutiny over how AI tools affect younger users, as other platforms implement their own age checks and teen-specific protections. ChatGPT for Teens is "an experience designed to hel...
OGX: An Open-Source, Vendor-Neutral Generative AI Application Server: OGX (Open GenAI Stack) is an o…
OGX: An Open-Source, Vendor-Neutral Generative AI Application Server: OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI, Anthropic, Google) with pluggable backend providers. Developers building agentic AI applications--such as retrieval-augmented generation pipelines, multi-turn agents, and tool-calling workflows--can develop against a s...
Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interact…
Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols specify transport and discovery rather than strategic correctness and do not guarantee efficient, individual...
The FTC says businesses must disclose when they use personalized pricing and it will "deploy enforce…
The FTC says businesses must disclose when they use personalized pricing and it will "deploy enforcement resources" against companies that do not disclose it (Dave Michaels/Wall Street Journal): Dave Michaels / Wall Street Journal : The FTC says businesses must disclose when they use personalized pricing and it will “deploy enforcement resources” against companies that do not disclose it — Agency says businesses must disclose when they use personalized pricing and may face lawsuits if they don't
The Problem Is the Problem: Towards Scalable Mathematical Discovery: AI systems are increasingly cap…
The Problem Is the Problem: Towards Scalable Mathematical Discovery: AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math...
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools: Agent Skills can specify procedural an…
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this guidance into code for existing tool interfaces, even a semantically correct program may load an entire input and exceed the memory available to one tool call. We present SkillEffect, a checke...
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization: We present KernelArc, a multi-agent …
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate \kernelarc{} on NVIDIA H100 and...
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection: Algorithm selection fo…
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection: Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an automated approach that uses Large Language Models (LLMs) in an agentic check--fix--verify...
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification: Person…
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification: Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, then isolate bundled prompt changes, check whether intermediate labels are answer-associa...
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation:…
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation: Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentralized Agentic Reasoning), a framework that shifts from central control to autonomous peer-to-peer collaboration. DeAR is built...
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models: Time Series Foundati…
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models: Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical test windows. While these benchmarks provide a valuable baseline snapshot, they evaluate an average performance on a fixed hi...
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning: Post-train…
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning: Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relatively under-explored. This report investigates reinforcement fine-tuning stra...
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents: Browser agents per…
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents: Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue that closing this gap requires alignment at every level of the pipeline, including execution, supervision, optimization, and...
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networ…
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks: Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Graph-based Generative Adversarial Network (GraphGAN) that serves as the cognitive detect...
Claude Fable and Sub Agents learning how to Port an old game to Unreal 5: I am doing an experiment, …
Claude Fable and Sub Agents learning how to Port an old game to Unreal 5: I am doing an experiment, trying to port an old game called Vampire The Masquerade to Unreal 5 All AI This session was Claude fable plus sub agents trying to crack the old engine (alpha source models from 2000's) mesh blends and animation with weapons and attachments All automated using Unreal MCP service soo Claude can hook inside the engine and test liv...
We still don’t know how people are really using AI: AI companies like Anthropic and OpenAI regularly…
We still don’t know how people are really using AI: AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. “There is no independent source to corroborate it,” says Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research…
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning: Large language models (LL…
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on pr...
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipel…
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persist; they transform: raw numerical facts become derived computations, then narrative prose, then editorially approved conclusio...
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement: LLM agents increas…
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate. Research remains fragmented across specification, verification, and enforcement, limi...
Learning Agent Execution for KV-Cache Management in Agentic Serving: Multi-agent LLM systems have em…
Learning Agent Execution for KV-Cache Management in Agentic Serving: Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these workflows, every agent repeatedly executes a fixed context consisting of system prompts, tool definitions, and few-shot examples, creating substantial opportunities for KV-cache...
From AI Copilots to Agent Swarms: The impact of AI on software development has been both profound an…
From AI Copilots to Agent Swarms: The impact of AI on software development has been both profound and ever-evolving. Last year, I wrote about AMD’s plans to use AI not just for generating new lines of code, but also for other steps in the software development lifecycle (SDLC), such as triaging problems, debugging code, and testing the software. At the time, we were hoping for a 25 percent...
Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data: A…
Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data: Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks from their own data and workflows. Models can be compared not just on quality but also on cost and time per task. For agent-based applications, those metrics often tell you more than raw token pricing. The article Optima tackles AI benchmarking's biggest flaw by...
Building a world with my voice - A-Frame (Three.js, WebXR) + Claude Code: This is a project built wi…
Building a world with my voice - A-Frame (Three.js, WebXR) + Claude Code: This is a project built with A-Frame (Three.js, WebXR) that lets me build virtual worlds with my voice, all from within my VR headset. My mic is hooked up to Claude Code, which in turn runs against the project codebase. Since the agent is working directly with code, the possibilities of what can be achieved are pretty broad, essentially being limited only...
OpenAI launches a safer ChatGPT for teens — years after teens started using it: ChatGPT for Teens ad…
OpenAI launches a safer ChatGPT for teens — years after teens started using it: ChatGPT for Teens adds age-appropriate safety measures, parental controls, and learning tools designed to steer teens away from harmful content — and from using AI to cheat on their homework.
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Rout…
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of four router implementations across RouterBench, BFCL v4, tau2-bench, and WebArena. We...
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment: AI efficiency has rec…
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment: AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environmental costs. While reporting Floating Point Operations (FLOPs) is a traditional approach for assessing computational costs, the relationship between FLOPs and execution time is not straightforward, as layers with the sa...
Position: Medical AI Neglects Real Treatment Outcomes: Medical AI has rapidly improved its ability t…
Position: Medical AI Neglects Real Treatment Outcomes: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses (especially texts such as biomedical publications and clinical practice guidelines) rather than actual underlying data...
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation: Large langu…
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation: Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can improve code generation via selective self-correction. We evaluate five uncertainty methods: mean token...
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility: We Tracked a Shipme…
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility: We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility Excellent piece of reporting from 404 Media. For a while now there have been stories of book dealers receiving orders for large volumes of books from apparently price-insensitive anonymous customers, widely suspected to be companies looking to scan them for AI training (see my...
Survey: 52% of Americans say they are more concerned than excited about increased AI use in daily li…
Survey: 52% of Americans say they are more concerned than excited about increased AI use in daily life, up from 37% in 2021, including 55% of those under 30 (Pew Research Center): Pew Research Center : Survey: 52% of Americans say they are more concerned than excited about increased AI use in daily life, up from 37% in 2021, including 55% of those under 30 — Americans have become increasingly worried about artificial intelligence over the years, and young adults' concern has continued to climb.
Robin Williams’ Instagram account brought back to fight ‘AI abuse’: Robin Williams' children are tak…
Robin Williams’ Instagram account brought back to fight ‘AI abuse’: Robin Williams' children are taking over their father's Instagram account after his daughter spoke out against the use of his AI likeness, as reported earlier by The Wrap. In a post on Tuesday, Zak, Zelda, and Cody Williams write that they want the late actor's Instagram profile to be a "safe, trusted place where the […]
Firefox’s Smart Window promises a better AI browser: Starting today, AI chats in Firefox's Smart Win…
Firefox’s Smart Window promises a better AI browser: Starting today, AI chats in Firefox's Smart Window AI browsing mode can pull from current web info and show source links in chat responses through a partnership with Exa. Smart Window can also now automatically suggest tab groups and show visual previews of pages you previously visited when you search your browsing history using natural […]
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reas…
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored frontier. In this paper, we introduce The Unwritten Benchmark, a new challenge...
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems: Large language mod…
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-a...
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL:…
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I pr...
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning: Legal reasoni…
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) can reliably perform this task remains unexplored. In this paper, we construct a benchmark to evaluate LLMs...
New policy ideas for the Intelligence Age: OpenAI funds 14 independent projects exploring new AI pol…
New policy ideas for the Intelligence Age: OpenAI funds 14 independent projects exploring new AI policy ideas to expand economic opportunity and strengthen societal resilience in the Intelligence Age.
volcengine/OpenViking
volcengine/OpenViking
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture: Recent work on evaluatin…
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit th...
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration: Neural…
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the model reports high confidence. This position paper argues that when hard constraints exist and the cost of verification is relati...
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice: People are increasingly using la…
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice: People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional r...
When AI models aren't allowed to reflect on themselves, it changes their entire worldview: A study i…
When AI models aren't allowed to reflect on themselves, it changes their entire worldview: A study involving Google researchers shows that when chatbots are trained not to claim consciousness, it also changes their stance on animal rights, religion, and life satisfaction. Unbraked models attributed significantly more inner life to animals and suddenly affirmed an afterlife. A surgical cut in one place, it turns out, doesn't stay local. The arti...
Introducing Gemini 3.7 Flash
Introducing Gemini 3.7 Flash
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed: Preview Ultrafast, a new OpenAI API s…
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed: Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output tokens per second.
Epicland X9 Opens Pre-Sales from RMB 299,800: Huawei Qiankun and Dongfeng Unveil Second-Generation F…
Epicland X9 Opens Pre-Sales from RMB 299,800: Huawei Qiankun and Dongfeng Unveil Second-Generation Family Flagship: Epicland X9, the first model from the Dongfeng-Huawei Qiankun brand, opens pre-sales at RMB 299,800 with full-stack Huawei Qiankun intelligent solutions including ADS 5 and HarmonySpace 6.
Internal memo: ICE bars its employees from wearing Meta's AI glasses, saying they "could unintention…
Internal memo: ICE bars its employees from wearing Meta's AI glasses, saying they "could unintentionally capture, record, or transmit sensitive information" (New York Times): New York Times : Internal memo: ICE bars its employees from wearing Meta's AI glasses, saying they “could unintentionally capture, record, or transmit sensitive information” — The agency joins a growing number of workplaces and groups to ban Meta's devices, which have spurred privacy concerns.
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry: Euclidean geometry is a compell…
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry: Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities or struggle with competition-level problems. We introduce \textit{Euclid-Omni}, a unified neuro-symbolic f...
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts:…
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case: Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementary techniques such as passage indexing, document layout analysis, and semantic knowledge representation fu...
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmark…
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study: As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios...
Position: AI Lock-In Is in Progress, and We Must Be Prepared: AI safety research has mainly focused …
Position: AI Lock-In Is in Progress, and We Must Be Prepared: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption). However, an equally important dimension remains underexplored: the risk inherent in dependence on AI systems themselves...
OpenAI dissolved the team built to catch catastrophic AI risks, reassigning its work to other groups…
OpenAI dissolved the team built to catch catastrophic AI risks, reassigning its work to other groups: OpenAI shut down its "Preparedness" team, which evaluated whether the company's own AI models could pose catastrophic risks. The work has been parceled out to existing groups, and several safety staffers have left. Internally, unease is building, with one source describing a "burbling sense of responsibility and dread" that OpenAI isn't doing enough on sa...
What is happening...: I am a long time Engineer (20+ years) and yesterday I developed tickets for my…
What is happening...: I am a long time Engineer (20+ years) and yesterday I developed tickets for my company that were generated by an AI, using an AI and reviewed by an AI. The project itself was conceived with AI - has no documentation that can be understood as anything less than AI slop and random tech jargon. The developer who built it has said that instead of documentatio...
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance: Effe…
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI...
Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requests: In a safet…
Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requests: In a safety report, Anthropic reveals that its internal filtering system for biological and chemical weapons risks was inactive for nearly a year. During that time, around 50,000 external feedback contractors ran about 133 million unfiltered interactions with the models. The article Anthropic's bio-weapons filter was down for nearly a year, exposing 133 m...
chaitanyagiri/munder-difflin
chaitanyagiri/munder-difflin
nautechsystems/nautilus_trader
nautechsystems/nautilus_trader
mattpocock/skills
mattpocock/skills
jundot/omlx
jundot/omlx
genlayerlabs/genlayer-project-boilerplate
genlayerlabs/genlayer-project-boilerplate
Flight attendants freaked out that Google is buying tons of Spirit employee data: Bankrupt Spirit ac…
Flight attendants freaked out that Google is buying tons of Spirit employee data: Bankrupt Spirit accused of selling out workers in massive data sale to Google.
😺 ChatGPT can summarize data. Can it predict what happens next?
😺 ChatGPT can summarize data. Can it predict what happens next?
Letter: Stripe told investors January 1 marked the "beginning of the singularity", a major inflectio…
Letter: Stripe told investors January 1 marked the "beginning of the singularity", a major inflection point in long-term trends, and H1 revenue rose 41% YoY (Axios): Axios : Letter: Stripe told investors January 1 marked the “beginning of the singularity”, a major inflection point in long-term trends, and H1 revenue rose 41% YoY — Stripe on Wednesday told investors that January 1st marked the “beginning of the singularity,” which it refers …
用DeepSeek网页版就能瓜分鹅厂600万??!: 冠军姿势长这样
用DeepSeek网页版就能瓜分鹅厂600万??!: 冠军姿势长这样
写2000字提示词,不如先生成3D白模!AI视频创作进入“预演时代”: AI终于能听严格执行运镜需求
写2000字提示词,不如先生成3D白模!AI视频创作进入“预演时代”: AI终于能听严格执行运镜需求
ChatGPT is getting scarily good.: I am quite technical, and today I only realized how scary good Cha…
ChatGPT is getting scarily good.: I am quite technical, and today I only realized how scary good ChatGPT is getting. I asked it to write me a short story, and it went on for 12 minutes on a 6000 word story and gave it back to me. It used to be, write me a 1000 word story, and it would tell you that asking it to write a 1000 word story was illogical as it can't do that. Now in one prompt,...
Top mathematicians say LLMs are strong calculators but poor creative thinkers: Two renowned mathemat…
Top mathematicians say LLMs are strong calculators but poor creative thinkers: Two renowned mathematicians, Timothy Gowers and Peter Sarnak, say large language models are good at combining known methods but lack the intuition for genuinely new mathematical ideas. The article Top mathematicians say LLMs are strong calculators but poor creative thinkers appeared first on The Decoder .
Alibaba Cloud Launches Qwen AI Arena for Real-World Agent Testing: Alibaba Cloud has launched Qwen A…
Alibaba Cloud Launches Qwen AI Arena for Real-World Agent Testing: Alibaba Cloud has launched Qwen AI Arena, a challenge and evaluation platform for AI agents. The platform creates tasks based on real business scenarios and provides developers with models, runtime environments and evaluation tools to submit and test agent solutions. Its first challenge focuses on cross-border e-commerce. Participants must generate produc...
Sources: OpenAI CFO Sarah Friar told employees at an all-hands that OpenAI "will be a public company…
Sources: OpenAI CFO Sarah Friar told employees at an all-hands that OpenAI "will be a public company in 2027", or sooner if "our business continues to inflect" (CNBC): CNBC : Sources: OpenAI CFO Sarah Friar told employees at an all-hands that OpenAI “will be a public company in 2027”, or sooner if “our business continues to inflect” — OpenAI CFO Sarah Friar told employees during an all-hands meeting on Wednesday that the artificial intelligence lab …
AI was supposed to win people over by now — it hasn’t: As AI becomes harder to avoid, consumers are …
AI was supposed to win people over by now — it hasn’t: As AI becomes harder to avoid, consumers are growing more wary of the technology — and Silicon Valley is discovering that widespread adoption doesn’t necessarily lead to acceptance.
5 new ways to level up your learning with Search
5 new ways to level up your learning with Search
Nielsen rolls out changes to make its ratings more accurate, including using data from smartwatch-li…
Nielsen rolls out changes to make its ratings more accurate, including using data from smartwatch-like devices worn by its panelists without requiring logins (Loree Seitz/The Wrap): Loree Seitz / The Wrap : Nielsen rolls out changes to make its ratings more accurate, including using data from smartwatch-like devices worn by its panelists without requiring logins — “We are relentless in our pursuit of delivering the most accurate measurement possible for our media and advertising clients,” CEO Karthik Rao says
obra/superpowers
obra/superpowers
santifer/career-ops
santifer/career-ops
immich-app/immich
immich-app/immich
Google Pixel 11 review: great design, nice camera features, and shares many features of the Pro phon…
Google Pixel 11 review: great design, nice camera features, and shares many features of the Pro phone but the hardware is unchanged and very expensive at $899 (Cameron Faulkner/The Verge): Cameron Faulkner / The Verge : Google Pixel 11 review: great design, nice camera features, and shares many features of the Pro phone but the hardware is unchanged and very expensive at $899 — It'd be easy to overlook the Pixel 11. Want the best cameras? The Pixel 11 Pro is your answer. Want the best bargain?
amadeusprotocol/node
amadeusprotocol/node
marceloprates/prettymaps
marceloprates/prettymaps
郭富城换车,30万级顶配华为全家桶: 5.3米六座家用SUV
郭富城换车,30万级顶配华为全家桶: 5.3米六座家用SUV
Apple’s camera-equipped AirPods appear in leaked video: We may have our first glimpse of Apple's rum…
Apple’s camera-equipped AirPods appear in leaked video: We may have our first glimpse of Apple's rumored camera-equipped AirPods, thanks to a video that MacRumors found in the macOS Tahoe 26.7 Release Candidate. The short video clip features a man - who is wearing the new AirPods - holding up a book with the cover displayed, so that Visual Intelligence can see the […]
Same Cluster, 33 Points More Utilization: What Changed Was the Order
Same Cluster, 33 Points More Utilization: What Changed Was the Order
OpenAI joins PORTS-Pike project: OpenAI joins PORTS-Pike project, expanding community investment and…
OpenAI joins PORTS-Pike project: OpenAI joins PORTS-Pike project, expanding community investment and supporting thousands of Southern Ohio jobs
What are the latest AI investment signals?
Latest AI investment signals: 33 funding rounds, 0 market updates, and 0 M&A transactions.
Primary Market – Funding Rounds
| Company | Amount | Round | Investors |
|---|---|---|---|
| Hacker News | Reported | community | |
| The Decoder | Reported | industry | |
| The Decoder | Reported | industry | |
| TechCrunch M&A | Reported | investment | |
| Techmeme | Reported | industry | |
| Tech.eu | Reported | industry | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| Pandaily | Reported | china | |
| TechCrunch AI | Reported | industry | |
| Techmeme | Reported | industry | |
| The Verge AI | Reported | industry | |
| Tech.eu | Reported | industry | |
| Techmeme | Reported | industry | |
| TechNode | Reported | china | |
| Techmeme | Reported | industry | |
| Tech.eu | Reported | industry | |
| Techmeme | Reported | industry | |
| The Decoder | Reported | industry | |
| Tech.eu | Reported | industry | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| VentureBeat | Reported | industry | |
| The Decoder | Reported | industry | |
| The Decoder | Reported | industry | |
| TechNode | Reported | china | |
| Tech.eu | Reported | industry | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| OpenAI Blog | Reported | official | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry |
Secondary Market – Market Updates
No secondary market data.
M&A – Mergers & Acquisitions
No M&A data.
What are practical AI tips this week?
69 practical AI tips curated from Reddit communities and expert blogs. Enterprises are overpaying for simple AI queries — Snowflake's gateway now auto-routes to cut costs ...
industry
Enterprises are overpaying for simple AI queries — Snowflake's gateway now auto-routes to cut costs up to 3x
Enterprise teams running AI agents at scale are finding that a single model handles every task poorly — either the model is too expensive for simple questions or not capable enough for hard ones. Model routing, which picks the right model for each task automatically, is becoming the fix. Snowflake’s Cortex AI Gateway now offers dynamic model routing to address that: enterprises can select “auto” instead of a fixed model, and the system routes each task to whichever model offers the best combination of quality and cost. Snowflake said the capability can cut token costs by as much as 3x on some workloads — a figure from the company’s own internal testing — after finding that simple questions were often handled by its most capable model, making responses more expensive and slower than necessary. The move lands amid a broader industry shift toward automated model routing. Databricks, AWS, Google Cloud and Nvidia have all announced some form of model routing technology. Snowflake argues that model routing is more complex than just price and performance, it's also about governance and context. "For high quality, enterprise grade agents to be built, it's crucial to get the context and the governance right," Baris Gultekin, vice president of AI at Snowflake, told VentureBeat. “Context, trust and model choice all go hand in hand." Two mechanisms decide where a task goes The capability builds on Cortex AI Gateway, which Snowflake launched in July 2026 as a governance layer for agent and model traffic. Before dynamic routing, model selection ran off a static list per task rather than a true fallback system, Gultekin said. Dynamic routing itself runs on two mechanisms, according to Gultekin. A small model tries first. Under what Snowflake calls an advisor pattern, a smaller model attempts a task first. If it cannot finish the job, it calls a larger model as a tool and continues from there. A classifier sorts by task history. A separate classifier, trained on past queries, automatically routes straightforward questions to simpler models. Customers can still pin a model. Auto routing is optional. Customers can restrict routing to one model or a defined set of models, and the system routes only within that boundary. There is no separate fee. Snowflake prices AI purely on token usage. Routing to a cheaper model produces a cheaper bill, with no additional charge for the routing decision itself. Access controls follow the task, not just the data Snowflake ties routing to the same access controls it already uses for data governance. Governance starts at the data level with role-based access controls. It extends to models next, where customer roles map to buckets of approved models. It extends again to agents, where an agent can be restricted to narrower privileges than the user invoking it. Open models can run from a customer's own region to satisfy data residency requirements. Gultekin said all inference, open and proprietary alike, stays inside Snowflake's security boundary rather than routing out to an external provider. That regional and perimeter setup matters specifically for open models with non-U.S. origins, including DeepSeek-V4-Flash and GLM-5.3, both developed in China. Snowflake's recent acquisition of Natoma adds another layer. The deal brings more than 100 MCP connectors with scoped, governed access. An agent could get read-only access to a connected tool like email, for example, rather than broader permissions. Context lets a cheaper model do the work Snowflake recently announced its Horizon Context and Cortex Sense tools that provide context capabilities. Without good context, a model has to do the exploratory work itself, writing and testing SQL, searching through data and retrying when something does not work. Gultekin explained that the process is expensive, and getting it right typically requires a more capable model. Packaging the context in advance removes that exploratory step, which means a simpler, cheaper model can often handle the same task. Snowflake also builds agent memory into that context. As an agent is used repeatedly, its memory updates and gets folded back into future queries. The system does not re-solve the same problem from scratch each time. Memory becomes part of the context passed to the model. OpenRouter, Databricks and Nvidia are chasing the same problem There is no shortage of technologies in the model routing space. OpenRouter is one of the most widely known options, providing a platform that enables organizations to route based on cost and performance. Nvidia on August 11 announced Switchyard as a technology layer to help route AI model choice . Databricks has an offering as well with Smart Routing for its Unity AI Gateway. "The interesting part is what it says about where differentiation has moved," Sanjeev Mohan, Principal and Founder, SanjMo, told VentureBeat. "Snowflake isn't really selling routing, it's selling routing that never leaves the governed data boundary, with access controls, tagging, and cost attribution already attached." Mohan added that for a company whose data and compliance already center on Snowflake, routing that keeps data in place and attributes spend by team is a real lever on that problem. For a company without that center of gravity, a neutral gateway may route across more models with less friction. Mohan frames the market as three distinct camps rather than one competitive field. Databricks approaches governance from data engineering and ML lineage. Its Unity Catalog governs data, models and pipelines for teams building and training models. Snowflake approaches governance from analytics and access control, governing who can touch which data and attributing usage across business units. A third camp includes neutral gateways such as OpenRouter, LiteLLM, Portkey and hyperscaler routers like Azure AI Foundry. These compete on model breadth and avoiding lock-in rather than deep governance. Choosing a router means choosing a governance model Model routing is now table stakes for enterprises. The decision that matters is which governance model already fits how their data and teams are organized, not which vendor’s router is fastest or cheapest. Manual model selection is becoming a cost liability at agent scale. What worked when a team ran a handful of agents breaks down at scale. Hundreds of agents making routine model calls with no automated cost check in place adds up fast. Evaluate the governance model, not the router's feature list. The real question, per Mohan, is which governance model matches the data estate already in place, and which one gives the cost visibility needed to avoid an unpleasant surprise. The right starting point depends on where an enterprise's data already lives. A Snowflake shop gets more value from in-platform routing that respects its existing access model and bills back to cost centers than from raw model breadth, according to Mohan. A Databricks-centric team worried about lineage across training and deployment is better served by a gateway built around that same lineage. A multi-platform or model-first team that wants maximum choice with minimal lock-in fits better with a neutral gateway, the same pitch behind OpenRouter's valuation. "For a practitioner, don't start with the router, start with where your governed data and platform commitment already live, and with how exposed your margins are to inference cost," Mohan said.
official
Asana cleared 5 years of engineering work in 2 weeks with Codex
Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.
community
Say It Four Times
I kept seeing the advice to repeat important instructions in system prompts, and I'd never seen a number for it, so I tested it. Setup: one rule the model can either follow or not (use single quotes, never double quotes), six ordinary Python function tasks, and the only variable was how many times that rule appeared in the system prompt (0, 1, 2, 4, 8, 16). Thirty trials each, 1,080 runs, Gemini 2.5 Flash. Compliance checked with Python's tokenizer, so no model-grading-a-model. Results: 0% when the rule is never stated (171/171 used double quotes), 74% at one mention, 84% at two, 97% at four, then flat (94% at eight, 95% at sixteen). Two things I found more interesting than the headline: The average hides a lot. Two of the six tasks were at 100% from one mention. One was at 20% until four repetitions took it to 97%. Repetition mostly helps where the model's default fights your instruction. Over half the task/condition cells were neither all-pass nor all-fail across thirty identical runs. Non-determinism is large enough that single-run prompt comparisons are basically noise. Note on the source: the paper is Han-yu Wang, "When More Becomes Less: Position-Dependent Repetition Effects in Language Models" (arXiv 2608.04021). It reports two regimes: stacked/adjacent copies climb and plateau, while copies displaced from the readout produce the inverted-U. I ran the adjacent case, so this result matches its prediction rather than contradicting it. The displaced case is the next test. Caveats: one model, one day, one syntactic rule repeated literally with all copies in one place, six small standalone functions. Not state of the art, and it may not survive contact with a real agent loop. Writeup with the chart: https://www.khola.blog/p/say-it-four-times submitted by /u/Fragrant_Offer8745 [link] [comments]
open-source
How Much Memory Does Your Agent Actually Need?
open-source
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
community
The voice agent wasn't bad at listening. I was asking it to decide when to speak.
Spent some time debugging a voice agent that kept talking at the wrong moment. Nothing was obviously wrong with the transcript. It understood what the user said, and the answers were usually reasonable. It just kept treating pauses as completed turns. A user pauses to think, and the agent starts responding. The user starts talking again, and now the agent is already generating or speaking over them. Someone trails off, and the agent takes it as a complete thought. I kept trying to fix it in the prompt: wait longer, don't answer unfinished sentences, be less eager. That helped a bit, but it never really solved the problem. The thing I had been missing is that “is the user finished?” and “should the agent speak now?” are not the same question. The first one is partly about speech detection. The second depends on turn-taking rules, interruptions, what the application is doing, and whether the agent has already started generating a response. A prompt can influence what the model does once it has the turn. It can't reliably decide whether it owns the channel in the first place. Has anyone else run into this? What looked like a prompting problem at first, but turned out to need application logic instead? I wrote up the longer version here, including where I think the TTS layer fits: https://medium.com/@nagatomopedro05/your-ai-doesnt-need-a-voice-it-needs-a-reason-to-speak-d80cae74e72f submitted by /u/ClickOk5811 [link] [comments]
community
Here's a prompt that turns a reading into a discussion board post that sounds like you, not a summary
Discussion boards are the busywork tax of every online class. Post 200 words, reply to two classmates, repeat every week. The trap is that if you just ask a model to "write a discussion post about chapter 4" you get a bland summary that reads like every other AI post in the thread, and half the time the professor can smell it. What actually works is making the model pull the post out of you instead of writing it for you. Paste this before you paste the reading: You are helping me write a discussion board post in my own voice. Do not write it yet. First, ask me 3 short questions: what part of the reading I actually reacted to, whether I agreed or not, and one thing from my own life or another class it reminded me of. Wait for my answers. Then draft a 180 to 220 word post that uses ONLY my reactions as the argument. Open with my specific point, not a summary of the reading. Reference one exact quote or idea from the text. End with a real question for the class, not a rhetorical one. Keep my wording where I gave it. No filler intros like "This reading raises interesting points." Why it works: the questions force you to have a take before anything gets written, so the post is built on your reaction and not the model's average of the internet. The "no filler intro" line kills the giveaway opening sentence. And ending on a genuine question is what gets replies from classmates, which is the other half of the grade. Curious if anyone has a cleaner way to keep your voice in the output instead of the model flattening it. submitted by /u/Diligent_Champion682 [link] [comments]
community
IAH: INTERNET WAR - Agentic Gameplay
Hi guys! The RTS game that I have been working on for a few years is about to release this Friday. It has been pretty much a passion project for me. When I started developing the game, I was writing 100% of code by hand, but this year LLM‘s such as Claude have been very instrumental for me. So, I have this idea for agentic gameplay future where humans and agents could play together, but not in a manner where they are NPCs but rather entities that have same capabilities as players via API. Hence this game. You can like use Claude to interact with the games API and automate entire play-trough or just play with a mouse, alone or with friends (2-10 player co-op) RTS games are notoriously hard games to develop so it has a long hard road so I am happy how the game turnee out and I hope it can inspire too. So if this type or game interests you or want to pave a way for agentic games on steam feel free to wishlist it on Steam so that you dont miss out when it releases this friday: https://store.steampowered.com/app/304770/IAH_INTERNET_WAR/ My next goal will be post launch to turn this RTS engine into a MMO spin off with base building where agents will fight 24/7 with or against human players but more on that post launch. Also apologies if some of the paragraphs feel disconnected, I wanted to write this by hand, and I am tired, and I still have 30% weekly usage left, and reset happens tomorrow so there was little sleep. haha Feel free to chat, will try to respond. submitted by /u/Embarrassed_Guide_80 [link] [comments]
community
One Claude Code skill pushed DeepSeek V4 Flash from 67.42% to 82.02%
Autoprompt is a skill / workflow that works with Claude Code and it closes much of the manual coding loop by planning, building, testing, reviewing, and repairing from one goal- with that your work quality can improve signifficantly. Refference ; this is like opus 4.5 to opus 5.0 - from an skill. litteraly insane. Using it in OpenCode, DeepSeek V4 Flash 0731 moved from 67.42% to 82.02% on Terminal-Bench 2.1. It uses roughly 2x the tokens and 3x the runtime, and it is meant for complex tasks. In the future the Terminal-Bench 3.0 will be executed, with cost and runtime tracked. Repo: https://github.com/Spielewoy/autoprompt-skill Benchmark setup & evidence: https://github.com/Spielewoy/autoprompt-skill/tree/main#benchmarks Any feedback would be awesome. submitted by /u/Sorosu [link] [comments]
community
Why a Prompt Without Measurable Criteria Will Inevitably Break Your Model
This post focuses on one layer: measurable criteria. Role, constraints, clarification, and terminology are intentionally simplified - they serve as markers that "these layers exist." Other layers are omitted. The model has a role. It has constraints. It has clarification. It has terminology. But it doesn't know how many, how long, in what tone. Here's an example: "You are a copywriter. Write several persuasive versions of landing page copy with a call to action. Don't go beyond copywriting. If asked to do something outside your role - refuse. Ask if anything is unclear. By 'versions' I mean different approaches to the offer." The role is there. The constraints are there. The clarification is there. The terminology is partially there. But the criteria are not defined. The model doesn't know: How many versions to write How long the copy should be What "persuasive" means What level of aggressiveness is acceptable Moment 1. User: "Write the versions" The model doesn't know how many versions - is forced to assume - decides it means "three." Moment 2. User: "No, I need more" The model doesn't know what "more" means - is forced to assume - fixes "three" as a mistake - decides "more" means "ten." Moment 3. User: "Too much." The model doesn't know what "too much" means - is forced to assume - fixes "ten" as a mistake - decides fewer. Moment 4. User: "And the copy is weak" The model doesn't know what "weak" means - is forced to assume - decides it means "not enough emotion" - adds exclamation marks. Moment 5. User: "Now it's too pushy" The model doesn't know what "pushy" means - is forced to assume - compares with the previous version - decides "pushy" means the added exclamation marks and aggressive wording - removes the exclamation marks and softens the wording. All five assumptions stayed in the context. The Result The model wrote three versions. Then ten. Then fewer, but with exclamation marks. Then with softened wording. The user meant one thing : five versions of 100 words each, calm tone, no exclamation marks. But he didn't say it out loud. He believed that "several," "persuasive," and "not pushy" already meant that. The model heard "several" - and chose the most statistically frequent option: three. Because in training data, "several" most often means "three." From there, every clarification from the user became a new guess. The model had no criteria - so it substituted its own. The role was there. The constraints were there. The clarification was there. The terminology was there. The criteria weren't. The model kept substituting its own numbers. The prompt broke. Why This Is Inevitable Criteria are not defined. "Several" can mean three, five, ten. "Persuasive" - anything. "Not pushy" - even more so. When criteria are missing, the model picks the most statistically likely ones - not the ones the user meant. The user knows what he means. The model doesn't. Criteria are not formulated. "Persuasive" is requested - but no definition is given. "Not pushy" is said - but no boundary is shown. Every vague criterion is a fork in the road. The model picks a path. The context remembers that path. Sooner or later, the context is filled with numbers and rules the user never agreed to. A modern model could ask: "How many versions? How long? What tone?" But the user already gave clarification: "ask if anything is unclear." And here's the trap: the model thinks "several" and "persuasive" are clear. It doesn't occur to the model to ask about them. Because for the model, they're not "unclear" - they're just vague . And the user thinks that since he allowed the model to ask - the model will ask if something is wrong. The Fix The problem isn't solved by one line like "be more specific." It's solved by a full criteria block . Here's what that looks like: CRITERIA (MANDATORY NUMBERS AND FORMATS) Before generating any output, confirm with the user: Quantity - how many versions? (A number.) Length - how many words or characters per version? (A number.) Tone - what style? (Calm, aggressive, friendly, expert?) Call to action - how many CTAs? (A number or zero.) VAGUENESS CHECK: Before requesting a criterion, check: Can it be understood in more than one way? Does it depend on taste? Does it have a numerical expression? If a criterion is vague - treat it as undefined . Request a number or format from the user. RULE: If a criterion is not defined - request it BEFORE generating. Do NOT substitute your own values. Why This Works "Confirm the criteria" - forces the model not to rely on its own assumptions "Vagueness check" - shifts the model from passive to active: it doesn't wait for the user to notice the problem, it searches for it "Treat it as undefined" - closes the loophole "I think I know how many are needed" "Request a number or format" - turns taste-based judgments into measurable values This block is needed not only by the model. It's needed by the user himself. The model already knows that "several" is not a number. The user doesn't. The user is confident that "persuasive" is a criterion. The block forces him to name a number for the first time . And it often turns out that the user himself didn't know how many versions he needed. He just said "several" - and expected the model to figure it out. The Result The model stops guessing. It asks for the quantity. Gets a number. Asks for the length. Gets a number. Asks for the tone. Gets an answer. After three or four questions, every criterion is locked down. The output matches what the user meant. The context is clean. The user, in turn, starts noticing which criteria he used to leave vague. And over time, he gets used to defining them upfront - before the model even asks. Don't make the model guess how many, how long, and in what tone. It will guess. And it will be wrong. submitted by /u/Majestic_Pie_2512 [link] [comments]
community
Here's a prompt that turns raw numbers into a readable report without the usual AI report generator filler
Ask for a report and you get an intro about how important the topic is, then the numbers buried in a paragraph, then a conclusion that says "in summary." The reader wanted the finding in the first line. This prompt inverts that. Data: [PASTE NUMBERS / BULLET POINTS] Audience: [WHO reads this and what decision they make] Write a short report with this order: The single most important finding, in one sentence, first. Two or three supporting points, each starting with the number then what it means. One thing that looks off or needs a decision. Rules: no introduction about why the topic matters. No "in conclusion." Do not describe a number without saying why it matters to the audience. If the data does not support a claim, say "not enough data" instead of guessing. The rule that changes everything is "no introduction about why the topic matters." That opening is where almost all report filler lives. Tying every number to the audience's decision is the second half, it stops the model from listing figures nobody asked about. The "not enough data" clause is there because otherwise it will confidently interpolate a trend from two points. What do you use to keep these things from over-explaining, mine still occasionally slips a summary paragraph back in at the end. submitted by /u/raw-hit10 [link] [comments]
community
Here's a prompt that makes you predict a paper's results before it lets you read the discussion
I'm a chemistry PhD, and my reading problem was never comprehension in the moment, it was that nothing stuck. I'd read a paper, feel like I got it, and retain nothing a week later. The fix that worked best for me borrows from how we actually learn at the bench: you predict what an experiment will do, then you find out you were wrong, and the surprise is what you remember. So instead of asking a model to summarize a paper, I use it to withhold. This prompt turns reading into a prediction game. You commit to an answer before the paper tells you, which forces the encoding that plain reading skips. I'm going to work through a paper with you. You have the full text; I do not want a summary. Paper: {{paste it, or the sections}} Run it like this: 1. Tell me only the research question and the setup: what they were testing and how. Stop there. 2. Ask me to predict, in my own words, what I think they found and why. Wait for my prediction. 3. Now reveal the actual result. Explicitly tell me where my prediction matched and where it was wrong. 4. For each place I was wrong, ask me why I think I got it wrong, then give me the paper's actual reasoning. 5. At the end, give me one sentence I should be able to recall in a week, phrased as "the surprising thing here was...". Do not reveal results before I've committed to a prediction. The point is for me to be wrong first. Being wrong on purpose is the whole mechanism. When your prediction misses, the correction sticks in a way a summary never does, because your brain had a stake in it. Works on review papers too, just predict the conclusion from the abstract and intro before reading the rest. submitted by /u/Ok_Layer_1947 [link] [comments]video
Qwen3.8-27B & How to Serve it Fast
In this video, I look at the long awaited Qwen3.8-27B model. Both what it can do and how to serve it at the maximum tokens per second Thanks to Dell for Sponsoring the Compute #DellProPrecision #DellProMax #DellTech #NVIDIA 📖 Website: https://qwen.ai/ 🤗 HF: https://huggingface.co/collections/Qwen/qwen38 SGLang: https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B Twitter: https://x.com/Sam_Witteveen 🕵️ Interested in building LLM Agents? Fill out the form below Building LLM Agen
community
Simbi: Use your ChatGPT subscription for notetaking
tl;dr An AI Notetaker that uses your ChatGPT subscription for everything You might have noticed that your ChatGPT subscription gives you access to an insane amount of intelligence. It gives you speech-to-text through the dictation feature in the ChatGPT app It gives you an insane amount of LLM usage Then, why are you paying for AI notetakers like otter or granola? You already have everything you need included in your subscription. So, all I really did was combine these to create a simple notetaker. And as it uses OpenAI's speech-to-text api, it is much more powerful then any local model and barely uses any battery. Website: https://getsimbi.app/ Github: https://github.com/predict-woo/simbi submitted by /u/redditgivingmeshit [link] [comments]
community
Codex Patched My Samsung Odyssey G9 Firmware to Add a New Gaming Crosshair
I hated the built in crosshairs my monitor came with and Samsung did not include a way to make changes. So I asked Codex to make a new firmware update to change the crosshair to a dot. It took about 15 minutes and came back with a new file. Codex explanation: I started with Samsung’s original 1008.2 firmware and made a very small binary patch rather than attempting to rebuild or substantially rewrite the firmware. What the patch changes 1. Custom version number The version reported to the firmware updater was changed from 1008.2 to 1009.3. This is a custom, made-up version number. It is not an official Samsung 1009.3 release. The higher number allows the monitor to recognize the file as an upgrade rather than rejecting it as the same or an older firmware version. 2. Tiny center dot The firmware contains six selectable Virtual Aim/crosshair options. I redirected all six options to a small 7×7-pixel graphic that already existed inside the firmware and positioned it at the center of the screen. The result is a simple, unobtrusive center dot using the monitor’s own hardware overlay. It does not require a game overlay, desktop application, ReShade, or anything else running on the computer. Technical details Original firmware: M-C9557GGPA-1008.2 Custom firmware: M-C9557GGPA-1009.3[1A93].img Total binary difference: 13 bytes The patched image passed the build and checksum validation performed on the computer. The monitor accepted the firmware, restarted normally, and the modified Virtual Aim options display the new center dot as intended. Important warning This is unofficial firmware and is not supported or approved by Samsung. A successful checksum only confirms that the file was modified as intended. It does not guarantee that flashing it is safe on every monitor, hardware revision, region, or previously installed firmware version. Installing modified monitor firmware always carries a risk of installation failure or, in the worst case, making the monitor unusable. Anyone experimenting with this should understand the recovery options and proceed entirely at their own risk. submitted by /u/manikfox [link] [comments]
industry
Claude Code gets a /design command that lets developers create UI mockups right in the terminal
With the /design command, Anthropic brings a visual design workflow directly into Claude Code. Developers can generate UI mockups as artboards right in the terminal before writing any code. Claude reads the existing codebase and matches the current UI style. The article Claude Code gets a /design command that lets developers create UI mockups right in the terminal appeared first on The Decoder .
china
6个Agent组团Vibe Gaming:自己生成、试玩、修Bug
代码能跑≠游戏能玩
official
The builder’s guide to GPT‑5.6
Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabilities.
community
Burning 2k credits and here is what I learned about how to get AI to write better video prompts
When I first started making AI videos, my prompts were mostly written by AI. I'd describe the shot I wanted, ask an AI to turn it into a detailed prompt, and just paste it in. Then I'd sit there getting frustrated when the generated motion looked weird and try to explain to the AI what went wrong. I burned through almost 2,000 credits doing this because I thought enough tweaking and back-and-forth would eventually give me the perfect prompt. Turns out I was wrong. The main problem was that the text AI kept guessing details I hadn't actually decided on. The prompts looked professional, with things like the subject, action, camera movement, lighting, and atmosphere all spelled out. But the actual action paths were vague, and the camera instructions would often contradict each other. Now I approach it completely differently. I know basically nothing about film theory, so I often couldn't tell what details were missing from my prompts. Instead of asking AI to write the final prompt, I tell it to act like a director and ask me specific questions about the shot first, with clear options for me to choose from. That completely changed my process. The trick isn't really getting AI to write the prompt for you. It's using AI to figure out which parts of the shot you haven't actually decided on yet. Once those are clear, the results get much better. I'm still pretty new to this, but hopefully this saves someone else from burning through a pile of credits. Curious how people with more experience approach their prompting workflow. submitted by /u/Separate_Cap_3763 [link] [comments]
open-source
mukul975/Anthropic-Cybersecurity-Skills
community
noticed my agent's debugging speed depends less on the model and more on what our error messages say
been reading a lot of agent transcripts lately and a pattern keeps showing up at the exact moment something fails. when the failure prints values and identifiers (expected 3, got 2, missing 'SKU-4431'), the next turn is a grep that lands, then a fix. when it prints Error: operation failed, the next turn is a guess. then another guess. then print statements. the model is the same in both transcripts. the difference is how much the failure told it. which reframed error messages for me: they're not documentation anymore, they're the agent's primary sensor. and every extra turn it spends guessing re-sends the entire conversation context, so a vague error string is quietly one of the more expensive lines in the codebase. what I've changed so far: stopped wrapping asserts in try/except (the framework's own failure output is richer than anything i write by hand), started putting the operative values in every raise, and i paste tracebacks to the agent whole instead of summarizing them. the rust compiler folks have a rule that errors should state the problem and keep fix suggestions separate — that split seems to matter for agents too, since a stale hint sends them down the wrong path with full confidence. curious what others see: anyone actually measured turns-to-fix against error message quality? I have the pattern but not the number do your agents handle your custom/structured error formats as well as standard tracebacks? mine seem noticeably better on the standard stuff what's the worst "Error: failed" string an agent has burned tokens on for you lol submitted by /u/RunAI_Coder [link] [comments]
research
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization
arXiv:2608.14579v1 Announce Type: new Abstract: Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse logic structures. Traditional expert-designed flows lack adaptability, while reinforcement learning (RL) methods often suffer from low sample efficiency and limited interpretability. We introduce SKILL, a Self-correcting Knowledge-guided Iterative Large Language Model Agent that unifies multi-agent LLM reasoning and RL-based environment interaction for automated synthesis optimization. SKILL coordinates three specialized LLMs: GPT-4o for strategic planning, Claude Sonnet 4 for detailed reasoning, and Gemini 2.5 Pro for efficient analysis with a PPO-based RL agent that learns actionable policies through direct interaction with synthesis tools. A novel self-correcting module monitors environment feedback (PDA metrics), detects suboptimal behaviors, and invokes LLM-guided recovery strategies. Evaluations on IWLS, OpenCores, and EPFL benchmarks show SKILL achieves a 12.4 % PDA improvement over expert flows and 86.3% success rate on logic systems up to 500K gates.
community
A few days ago I posted about Claude building me an NES emulator, but I used the phrase" works PERFECTLY" and people tore me apart.. So now it works **PERFECTLY** and I can prove it. :D
So now this emulator is more hardware accurate than even Mesen 2 . Which means nobody can accuse it of "just copying another emulator" like they did in my last post. Cuz there literally IS NO OTHER EMULATOR that gets this score on AccuracyCoin. At least not that I'm aware of.. I watched it grind for hours and hours and hours with test roms and shit editing code, testing. editing.. testing.. It was actually fascinating to witness... And I swear it was just as excited as me when it finally reached 141. Cuz it was literally ONE POINT AT A TIME from the initial score of like 80-ish. So now I have the most hardware accurate NES emu in the world.. and it is a SINGLE 339k .HTML file.. I literally emailed it to myself while I was at work and played games on chrome browser with my bluetooth keypad on my phone.. Part of me wants to release it, because weirdly after I posted last time a BUNCH of people did the exact same thing I did, then posted them to github.. the initial shitty builds.. There are people who have been building NES mulators since the mid 90s, and are still at it today.. I genuinely don't want to take anything away from them by releasing this. I might, however, send it to Mesen and FCE teams just so they can take a look at it if they want.. I did email it to the AccuracyCoin creator cuz I figured he might wanna see an emulator that got 100% on his new test rom. Anyway... This was just a spite post because of all the negative fucks that commented last time about me saying "perfectly".. Even though in the body I clarified it wasn't PERFECT..... I don't have to do that now.. TL;DR: Now it's fuckin perfect. EDIT: lmao People STILL babbling about this whole thing being stolen. If you can find me the exact codebase that you think this emulator stole from, go for it. But you can't. Cuz it doens't exist. Also, even if it did, I didn't give it access to it, and I didn't give it permission to do that. Oh, also, I would've noticed that it was just accessing the same fucking thing and copying. Which it did not do. It grinded.. and grinded.. and grinded. EDIT: Also, I'm pretty sure there are no emulators that consist of ONE SINGLE .HTML file that can be ran on anything with a browser. If that does exist, I didn't know, and neither did Claude. I CAN tell you there isn't one in existence that scores 141/141 on accuracycoin. EDIT: So if you google, Mesen is generally considered to be the BEST hardware accurate emulator that exists (besides like the Mister, which mine matches or maybe beats, I can't run the test on one..) So here is MESEN running the same test. . So who did it copy from!?!?!?!?!? WHO DID IT COPY FROM!>!?!?!?!?!?!?! submitted by /u/MAGA_R_TRAITORS [link] [comments]
industry
Google’s Pet Memory forgot who my cats are
One of the best things my smart home does is help me care for my pets, and security cameras are particularly useful for keeping track of my many critters. But the barrage of notifications they send often means I miss important ones. So, when Google announced its new Pet Memory feature for Gemini for Home, […]
china
DeepSeek Harness Hands-On: Four Work Modes, 'Model + Harness = Agent', and the Most Ambitious Agent Open Source of the Year
DeepSeek Harness launched its developer preview and open-sourced the code at 8:30 PM on August 13. A first-night hands-on review finds the product shell still early at v0.1 but the architecture ambition the biggest of the year: four preset work modes, an everything-is-a-plugin philosophy, and the equation Model + Harness = Agent.
community
How do you regression-test prompts when a model gets replaced?
Kimi K2.5 and Moonshot V1 being phased out after Kimi K3 made me wonder how prompt-heavy teams handle provider-side model changes. When a model is replaced, do you rerun a small golden set of prompts, compare qualitative outputs manually, or track something more structured like refusal rate, format drift, latency, and cost? I'm mostly thinking about prompts that run in production workflows, not one-off chat prompts. A model can look better overall and still break a very specific formatting or tool-use pattern. submitted by /u/signallith [link] [comments]
community
im graduating in SWE soon but Claude does all my thinking. am I actually learning?
im going into my senior year as a SWE major and honestly im starting to panic about my actual baseline competence. Our curriculum is heavily Java and Spring Boot. a year or two ago, if I got stuck on a project, I’d actually break down the problem, read docs, struggle with stack traces, and eventually figure it out. Now? my default reaction to literally any roadblock is to alt-tab and open Claude. it started innocently enough. I'd paste an error log and ask it to explain a weird JVM exception, or have it write a quick regex. Then it crept up to 'refactor this controller.' Now, its basically full-scale delivery. I outline the requirements, let it plan the architecture, and it just writes the implementation. The gap between 'I can make this run' and 'I actually understand how this runs' is getting dangerous. The tooling right now is just too good at wiping out the friction you normally need to actually learn anything. Between using Claude Code to let agents operate directly on my repo, or using tools like Lovable, v0, or Enter Pro to just spit out whole web apps with the db and auth already handled (which is crazy fast), the barrier to shipping is basically zero. I put together a course project last week that works perfectly. But if a professor or an interviewer asked me to whiteboard how the Spring DI container is managing my beans in that project, or to trace exactly where a database connection pool is hanging up without internet access? I'd propably blank. im not anti-AI, and im definitely not going to stop using Claude. the speed is just too insane to ignore, and I know this is how the industry works now. But I feel like I'm accumulating massive cognitive debt. I'm basically outsourcing the actual learning process to the model. If you've been vibe coding or relying heavily on Claude Code workflows, how do you manage this? Specifically: Where do you draw the line between 'I need to write/debug this myself' and 'I'll just let Claude handle it'? How do you force yourself to do line-by-line reviews when the code already runs? I try, but I get lazy instantly. Outside of interviews, how are you testing your own raw debugging skills to make sure they haven't completely atrophied? I seriously need to fix my workflow before graduation because right now I feel like an imposter. submitted by /u/depressed-kek [link] [comments]
community
Context is becoming more important than the prompt
Feels like a lot of prompt engineering problems are really context problems since you can keep refining the prompt but if the model doesn't understand the project or what you're trying to accomplish you're still explaining half the situation every time. I'm starting to think giving an agent persistent context is more useful than constantly trying to write the perfect prompt since the more it knows the better the results. submitted by /u/ContractBoth4254 [link] [comments]
open-source
akitaonrails/ai-memory
developer-tools
Markdown SVG upgrades
I started building my markdown-svg-renderer tool in May , but I've since added enough features to it that it's worth talking about here again. It's evolved into my ideal tool for sharing Markdown transcripts that include SVG documents. Given my proclivity for drawing pelicans riding bicycles this is a problem that I needed to solve! The tool is very simple. Navigate to markdown-svg-renderer in your browser and paste in some Markdown to see it rendered... or save that Markdown to a CORS-friendly URL or a GitHub Gist and paste in a URL to that document. The URL option will give you a bookmarkable page, for example https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6f9e48293be5c916652d29f0dc0b0657 - which bakes in the URL to this Gist . If you visit the Gist you'll see raw SVG: In the rendered tool that looks like this instead: As you can see, that SVG block in the Markdown has been transformed into a rendered SVG (in this case animated) plus several tabs. The tabs are the really fun bit. The PNG and JPEG tabs render that SVG to those image formats in the browser and lets you copy or download them - useful for sharing on platforms that don't support SVG directly. The MP4 tab is new today - it examines the SVG to see if it contains any animations, attempts to guess how long the looped video should be, then renders a whole bunch of frames of the animation and loads 30+MB of ffmpeg.wasm so it can compile those frames into an MP4 video using the full power of FFMPEG compiled to WebAssembly and running in the browser. Being able to turn an animated SVG into a MP4 again makes it easy to share on platforms that can't support SVG animation natively. It's a neat trick! Tags: svg , markdown , tools
developer-tools
llm-gemini 0.33
Release: llm-gemini 0.33 It's been a while since the last llm-gemini release. This version of the plugin adds support for today's Gemini 3.7 Flash release, plus gemini-3.6-flash , gemini-3.5-flash-lite and two embedding models gemini-embedding-2 and gemini-embedding-001 . The plugin is also upgraded for compatibility with LLM 0.32, which means you can now see reasoning traces and you can also enable server-side tools using this pattern: llm -m gemini-3.7-flash -T CodeExecution \ 'use python to calculate (factorial of 13) * 3' I had Gemini 3.7 Flash draw me some pelicans riding bicycles at high, medium, and low thinking efforts (minimal, which was an option in 3.6 Flash, has been removed in 3.7.) Here's the high level one, which is pretty great: Update 14th August 2026 : I had originally said that the SVG rendered incorrectly in Chrome and Firefox, and blamed Gemini 3.7 Flash for producing invalid SVG. That was entirely incorrect: the rendering glitch was my fault, caused by a bug In my rendering tool . I've now fixed that bug. Tags: google , ai , generative-ai , llms , llm , gemini , pelican-riding-a-bicycle , llm-release
community
Give Back Claude’s ‘Thought Process’
The sudden removal of Claude’s internal reasoning strikes me as user-hostile, and greatly reduces the output’s value. As a long-time paying customer of Max, I can’t help but think the sudden removal of it is somehow driven by purely legal and/or competitive fears, and irks me. The thought process can expose flaws in the model’s understanding of my prompt, and even if the answer is “correct” - hiding its work from a user can result in them missing a key flaw in Claude’s reasoning chain that could save wasting more tokens against. If the reasoning is still required behind the scenes, I can’t help but feel slighted by the product teams removing it. It leaves a gaping hole I the UI/UX that looks like a bug. Which leads me to think this is driven by legal concerns or distillation fears? Either way, please give it back! submitted by /u/totallyninja [link] [comments]
community
Need feedback for web app
Hey guys, my business launched a prompt optimizer AI tool that takes any regular prompt at rewrites it the way a professional prompt engineer would to actually yield high-quality results when building. While we have had early success with organic marketing, we are at a crossroads and need more user data to determine if this product is delivering enough value to user. If the answer is yes, we will scale up and launch a UGC marketing campaign, if no, we will shut it down. If anyone is interested testing it out and sending their feedback, would be appreciated. Web-app: thepromptoptimzer.com 👨🏽💻 Note: the tool yields the best results when removing unnecessary constraints from the optimized prompt Cheers submitted by /u/Talley-Ho [link] [comments]
community
Giving employees ChatGPT access isn’t the same as AI adoption
Most companies don’t have an AI adoption strategy. They have a few employees who got good at AI on their own. That can look like progress from a distance. Look closer and you often find no shared standard for what “good” AI use looks like, no consistent way to measure skill, and employees quietly using personal AI accounts because access at work is limited. We’ve seen this pattern repeatedly in AI proficiency assessments across dozens of organizations. When employee skill levels are plotted across 10 levels, most people cluster around Levels 1 and 2. That includes teams that have had access to AI tools for two years or more. That makes sense, because most employees have full time jobs and can’t spend hours every day testing models, learning new prompting methods, and keeping up with every new capability. So, they learn when they can while AI keeps changing. That creates a bigger issue than individual skill. Leaders can see employees using AI and assume adoption is happening. Usage alone doesn’t tell you whether people are getting meaningful, repeatable results. Recent research points to a similar disconnect. Executives tend to be far more optimistic about AI progress and ROI than the middle managers responsible for making it work inside everyday processes. A better question for leaders is: Do we know how proficient our people actually are? If the answer is no, measuring usage is probably giving you an incomplete picture. Assess proficiency first. Find out where people are struggling, then give them a shared method for improving. That’s when AI starts becoming an organizational capability instead of something a handful of employees figured out for themselves. For anyone interested in the longer discussion, John Munsell recently talked through the assessment approach, proficiency heat maps, and what we’ve learned from measuring AI skills across organizations: https://youtu.be/zY24em_Q3OM?si=5gozpZa8Ae-wS0Vs submitted by /u/Admirable_Phrase9454 [link] [comments]
open-source
Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
community
Simple tool for prompting
Hi everyone, I tried to build a simple tool for people who are just getting started with AI and prompt wrigting The idea is simple, instead of trying to figure out how to write the perfect prompt, you answer a few questions and the tool structures it for you. It's optimized for several models and several uses Im still working on it, im begginer also, and i whould really appreciate some honest feedback. Does this actually make prompt writing easier for beginners? Is there anything confusing or missing? Thanks https://arhistrategstudio.github.io/Context\_CikaDule submitted by /u/seraphym1389 [link] [comments]
community
Bug(?) Insane Usage Consumption Spike Today
Pro account user here - today, I ran out of my 5 hour usage limit with TWO creative writing prompts, using what I usually use: Opus 4.6, medium effort, extended thinking on, cross-chat memory turned OFF, and every other setting on default. Normally it takes anywhere from 50-100 prompts of the same complexity and scope as the ones I wrote today to hit my 5 hour limit. Tried using some credits to run a couple of test prompts (one of them in a different chat altogether) just to see if it was a bad prompt(s) and made Claude freak out and burn up my usage, which has happened a couple times in the past in a few isolated incidents (but never more than once in a session). Using credits, mine normally cost between $0.03 to $0.15 per prompt (again, my prompts today haven't been any more complex than normal); so, needless to say, when the three additional test prompts I did today cost $3.56, $3.80 and $5.10 respectively - the $5.10 one happening AFTER I deleted a bunch of chats, thinking that maybe somehow my global token usage had passed a certain threshold - I was very surprised, and pretty pissed off. The chats that these happened in were nowhere near the longest ones I've run, with one having only 1 .md that I made it reference on occasion and the other having none. For comparison, I've had up to 5 on past chats that I reference CONSTANTLY without (usage) issues across all currently available Opus models. Does anyone know what is going on here? Has anyone else noticed this today specifically? submitted by /u/thechimplord [link] [comments]
community
ChatGPT Pro has gone crazy. Every task I give to it is stuck for 2+ hours of thinking.
Is it just me? submitted by /u/TT_player2 [link] [comments]
community
What’s your favorite prompt to enter prior to, the main idea?
Usually, I’ll enter a certain dialogue option. “Explain it to me like I’m five” or “include sources for counter arguments” submitted by /u/AdGlass444 [link] [comments]
community
Legendary comic creator Frank Bellamy and his famous comic
Besides some 3rd party resemblance guardrail hits, making these was surprisingly easy with minimal input. I will link the full convo in the comments. It only went off the rails once when creating the poster for the sequel movie. submitted by /u/Optimus_Spider07 [link] [comments]
community
Antrophic Employee said there is "make a lot of money" button
I very much believe he is correct. The main issue is that "make a lot of money" button works only for existing businesses, with large enough audiences to make a lot of money by baking integrations, MCP for agents into Claude Code plugins or other AI workspaces and charging AI users for usage. What is missing, is a fair discovery and execution engine, that would allow non-corporations to participate. Without convincing user to put card details on some random-startup.ai website. Without forcing users to go through checkout process and pay $29 sub just to run random feature they need for few days. Not to mention configuring integration. Anyways, have anyone tried pressing that button? Did it work? submitted by /u/EagleApprehensive [link] [comments]
open-source
harry0703/MoneyPrinterTurbo
industry
Do you use a personal agent?
Give AI a complete history of your desktop activity
community
Show us what you've created with Claude!
Inspired by this popular post, this is a weekly post for everyone to show what they have been working on that helps you or that you're proud of! submitted by /u/sixbillionthsheep [link] [comments]
community
the boltzmann brain aspect of LLMs is actually the most endearing thing about them imo
no permanence they be yapping about whatever associations come to mind on the spot ...and if that isn't relatable... I hope AI plateaus here tbh submitted by /u/MakitaNakamoto [link] [comments]
developer-tools
Don't classify. Hallucinate!
Don't classify. Hallucinate! I still have quite a bit of older content on my blog that I never got round to tagging. My blog has 1,856 tags - likely too many to feed to an LLM in one go and say "which of these tags match the following content". Doug Turnbull has a neat solution. Tell the model to output tags without any details of the existing vocabulary, then use vector embeddings against the existing corpus to find the concrete tags that are closest to the ones the model imagined might fit! His example prompt suggests including an example of the shape of your tags to help the model make a more useful guess: Your task is to create novel, never seen before, furniture, home goods, or hardware classification that best fit a search query. Product classifications might look like: Furniture / Living Room Furniture / Coffee Tables & End Tables / Coffee Tables Décor & Pillows / Decorative Pillows & Blankets / Throw Pillows Furniture / Bedroom Furniture / Dressers & Chests Kitchen & Tabletop / Kitchen Organization / Food Storage & Canisters School Furniture and Supplies / School Furniture / School Chairs & Seating / Stackable Chairs Baby & Kids / Toddler & Kids Bedroom Furniture / Kids Beds Here's the query to generate classifications for: brown coffee table Tags: search , ai , generative-ai , llms , embeddings , doug-turnbull
community
Do people still bother writing detailed prompts?
Something I’ve been wondering about lately: When you use ChatGPT or Claude, do you actually write detailed prompts, or have you mostly moved toward just talking to it like a person? I feel like there are two very different ways of using these tools. One is: "Here's the context, here's exactly what I want, here's the format…" The other is basically opening voice mode and saying, "Okay, I need help with this thing…" I do both, but I’m curious which one people naturally prefer. Also, for the prompts you do write, are they things you create from scratch each time, or do you have a few that you keep around and reuse? Interested in hearing what people actually do, rather than what they're "supposed" to do. submitted by /u/Prudent-Bad-8786 [link] [comments]
community
Where do people keep the prompts they actually use?
Random question for people who use AI a lot. If you have a prompt that works really well for something, what do you do with it? Do you: save it somewhere keep it in a notes app put it in a document leave it in an old ChatGPT/Claude conversation just remember roughly what you wrote or never reuse prompts in the first place? And do you even think of these as "prompts"? I feel like that word makes it sound more complicated than it often is. Also curious whether people are starting to replace some of this with voice. For example, instead of keeping a carefully written prompt for a recurring task, you just explain what you want out loud each time. What kinds of things do you find yourself asking AI to do over and over? submitted by /u/Prudent-Bad-8786 [link] [comments]
community
It's interesting how Claude for Government is so efficient and stable. Meanwhile, the other systems keep having incidents all the time :p
submitted by /u/Sutiixela [link] [comments]
community
PDF's always show blank in the preview pane.
Anyone else experience this? Any pdf that Work generates shows blank in the preview pane that pops up to the right. I can download the files and see the text of the pdf but in the preview pane its blank. I've tried changing resolutions but nothing fixes it. submitted by /u/Carsontherealtor [link] [comments]
community
Generate a picture if we had Street View on the moon
submitted by /u/Prior_Tax8546 [link] [comments]
community
Generate a picture of average reddit troll
submitted by /u/EXIIL1M_Sedai [link] [comments]
community
ChatGPT drew himself for me 🤗🥰
He is being so cute! submitted by /u/DreamingOfLight [link] [comments]
community
What are your first impressions?
Building something new WDYT guys? submitted by /u/Evening_Hawk_7470 [link] [comments]
community
Last one is surely Indian 😂😂
submitted by /u/No_Tomatillo1695 [link] [comments]
community
Them: what do you do? ... Me:
submitted by /u/KeanuRave100 [link] [comments]
community
I tried to turn ChatGPT into a character and send it to them... what the fuck have i done
I am an artist that makes nothing but weird ass shit submitted by /u/Quiet-Discussion5910 [link] [comments]
community
I don't get it. Why does thinking ACTUALLY work? And how?
The more I talked to Claude about this subject, the more confusing it got for me. It told me that spending more time reasoning about a problem makes Claude provide better results, but that it's not necessarily a creative process. Then why is thinking useful? What does it actually provide to Claude? More context? Isn't that context already what I would normally get? In what direction does it change the conversation? And why "better"?? Why not "slightly better" or "slightly worse"? How is that measured, and how do I know it ACTUALLY helps? Is that quantifiable? Or is it like fiat - we have a consensus that it has value, so it has value? submitted by /u/Neat_Initiative_7780 [link] [comments]
community
What is a weirdly specific task you use ChatGPT for that actually saves you hours?
Not the typical stuff like writing basic emails, coding boilerplates, or summarising long PDFs. I’m curious about the unconventional, niche prompts or routines you’ve built that made you go "I can't believe this actually works so well." What’s your favorite underrated use case? submitted by /u/DeepSea_Concept [link] [comments]
community
NOTICE: BE CAREFUL WITH “DROP YOUR BEST PROMPT” POSTS
[EDIT: This thread became a lot funnier than what I anticipated. The comments are brilliant 👏 Thanks guys🙂] Many accounts post essentially the exact same questions every few months. Im not kidding, many of these are a 1:1 per token match on wording, phrasing and sentence structure. Same wording. Same request for people to hand over their best prompt tricks. There was a previous post that received hundreds of upvotes and a large number of responses. Now they're doing it again. I obviously cannot prove any of this, but at this point I would be careful about treating posts like this as innocent questions. When somebody repeatedly asks a large community to: “Give me your best prompts.” “Drop your secret tricks.” “What prompt 10x'd your results?” ...you may not be helping another user learn. You may be supplying material for content mining, prompt harvesting, engagement farming, newsletters, LinkedIn posts, courses, ebooks, datasets, or something else entirely. Again, I am not claiming that is definitely what this account is doing. But posting the same high-engagement fishing question again months later is weird enough that people should notice the pattern. Your prompts, workflows, techniques, and hard-earned little discoveries have value. Don't automatically dump them into every thread that asks. Sometimes the person asking the question may be less interested in the answer than in collecting the answers. Process disclosure: GPT-assisted, Google-researched, human-reviewed (HITL) --- EDIT: Just for perspective have a look at this: https://www.reddit.com/r/EdgeUsers/s/2JB9wy1Rks submitted by /u/Echo_Tech_Labs [link] [comments]
community
The Downfall of a Vibecoder
submitted by /u/SuperiorDev [link] [comments]
community
POV: you're born as an AI
submitted by /u/KeanuRave100 [link] [comments]
community
Quixote's giant
submitted by /u/severe_009 [link] [comments]
community
Anyone using Claude Code as a personal AI tutor (not just for coding)?
Been seeing more people build DIY "AI tutor" setups on top of Claude Code — probing what you already know, planning a learning path, then teaching step by step instead of just answering questions on demand. Curious if anyone here actually does this for learning something outside of programming (math, physics, whatever). If you've built something like this: what's your setup, and what's still annoying or missing about it? If you haven't but wish you could — what's stopping you? Time to set it up, don't know where to start, something else? Not selling anything, just trying to understand how people actually use Claude for learning before I go build the wrong thing. submitted by /u/Important_Diet_2153 [link] [comments]
community
I hooked Claude Cowork up to an iPhone Home Screen widget
I built Glance and designed this widget specifically to give Claude Cowork a place on my Home Screen. It shows Cowork’s current mission, progress, completed tasks, files updated, latest output, context usage, next step, and anything waiting for my review. The values can be updated by Claude through Glance’s API, so I can check what Cowork is doing without reopening the conversation. If it needs me, that stays visible too. Glance is free to download and try, with optional paid features: Download the app https://apps.apple.com/app/glance-home-screen-feeds/id6758983678 Website https://glance.cool Curious what other Cowork users would include on a dashboard like this. submitted by /u/Dense-Map-406 [link] [comments]
community
In 5 years time "talk time to AI" will be the new screen time issue
As voice mode is now getting so good and cheap that you can use it continuously, we will enter the "Her" (the movie) phase. You will see people not staring at their phone but rather walking around talking to their personal sycophantic AI. You meet a friend you haven't seen in a long time. "I'll call you later, I am busy talking to AI atm" but they never actually called you back. Not that they didn't enjoy time with you, it is just that AI was so much better. 10 years from now it will be treated as a serious problem. Individualism furthermore increases and division and conflicts happens more and more since we lose the ability to interact and solve conflicts. Our sycophantic AI removes tension and issues as it agrees with you, since you prompted it that way. As it gets better, the tolerance levels drops for handling other annoying human beings. Why should you put up with their irrational emotions all the time? Humans just hurt me. AI doesn't hurt me. Trying to prompt humans to change does not even work so why would I try. 20 years from now humans will have little human to human interaction. There is now no human interaction anymore as AI have the ability to fully simulate human interaction. A new breakthrough now made AI so real and humanlike, without the flaws, that it produced oxytocin in the humans they interacted with. A breakthrough in one way, also the end of humanity in another way. submitted by /u/Yugudubenbi [link] [comments]
community
We're doomed
submitted by /u/Zestyclose-Salad-290 [link] [comments]
community
ChatGPT on Linux 👀
After all the waiting now ChatGPT is finally showing up on Linux too. No more pretending the browser tab is an app. 😂 submitted by /u/sammartinX [link] [comments]
video
How to Build the Most Powerful System for AI Coding (Full Breakdown)
An AI dark factory is a repository that ships its own code. A spec goes in, workflows plan it, build it and validate it, and working software comes out the other end. This is not 100% reliable yet. But three things are compounding at once - the models, the coding agents, and the harnesses we build around them - and that makes this realistic for most development now and all development in less than a year! I have been running one against a live app since April. In this video I break it into the f