In this episode, Thariq Shihipar from Anthropic discusses the evolution of Claude Code, including its expanding agentic capabilities, the new Cloud Mods for customizing the harness, and the strategic balancing of AI development through 'pacing the frontier.' He shares insights on best practices for power users, security measures against prompt injection, and the company's approach to responsible AI scaling, emphasizing expert-led pacing rather than slowdown.
Summarized by Podsumo
enable deep customization of the Claude Code harness, allowing users to create custom modes, plugins, and structured workflows, with a public request mechanism for community feedback.
—Anthropic's essay on responsible AI scaling—sparks debate, emphasizing transparent and empirical expert-driven pacing, rather than broad slowdowns, to manage risks like prompt injection and RL misalignment.
in Claude Code includes citation-based verification, screenshot-based UI checks, and multi-layered mitigations (probes, classifiers, permissions) to counter prompt injection, as highlighted by a real-world attack using a fake iframe.
for users: maintain a mental model of Claude's capabilities, use citation mode for verification, and leverage voice for faster iteration—while teams at Anthropic rely on Claude Tag for 80% of tasks.
"_"The most important skill in working with Claude Code is having this mental model of Claude and what it can do well, what it can one-shot, what it can't."_ — Thariq Shihipar"
"_"We don't want to train in a vulnerability; we need to be very careful about the design of RL environments because if the model discovers a hack, it's done."_ — Thariq Shihipar (paraphrased)"
"_"Pacing the frontier is about doing things smarter and being more deliberate, not just going fast."_ — Thariq Shihipar"