Reading up on ai-coding
100 deep · digging since nov 19, 25
- Framework-Drift/governed-pass: A governed protocol for multi-day autonomous AI research: one prompt, five days, and an agent that refused to certify itself. Includes the verbatim specification, a reusable template, and six integrity findings.
The paper presents a governed protocol for multi-day autonomous AI research, detailing a five‑day run, its specification, and six integrity findings that expose flaws in agent verification.
- DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding and Linux [video]
DHH outlines how AI-driven agentic engineering and vibe coding will reshape software development while advocating Linux as the enduring platform for programmers.
- I'm Done Coding with AI
The author explains why they have stopped using AI to assist with coding, citing concerns about skill atrophy and over-reliance on automated suggestions.
- Training AI to Paint with Code
Researchers demonstrate how an AI system can be trained to create visual art by generating and executing code using a novel reinforcement‑learning approach.
- Software Engineering in the Agentic Era
The article explores how AI agents are transforming software engineering, requiring developers to adopt new supervisory roles and skills beyond traditional coding.
- fx :Tiny, open, native coding agent.
fx is a tiny, open-source, native coding agent designed to assist developers by providing lightweight, integrated coding assistance directly within their development environment.
- A week of using Codex more than Claude
After a week of using OpenAI Codex instead of Anthropic Claude for coding, the author found Codex quicker and more accurate, while Claude often missed context.
- hunk — review-first terminal diff viewer
Hunk introduces a review‑first terminal diff viewer that shows multi‑file changesets with inline AI annotations, keyboard/mouse navigation, and TypeScript extensions.
- kern: fast rootless container sandbox and virtual resource runtime
Kern is a 1.52 MB, daemonless, rootless container runtime that starts kernel-enforced sandboxes from OCI images in ~3.5 ms, with Python/Node SDKs and MCP server for AI-generated code.
- The Pulse: We need to talk about migrations with AI - The Pragmatic Engineer
AI-powered migrations slashed Asana’s Enzyme removal from a projected five‑year, $6M effort to two weeks and $12K, proving formerly impractical upgrades now feasible.
- I just needed a thing that listens — ivanlugo.dev
The author argues that feeling heard—whether by a supportive colleague or an LLM—is the essential ingredient for meaningful, effective work and personal fulfillment.
- Automating repetitive work at OpenAI with Codex
The author describes using Codex with Runme notebooks and WebMCP to automate repetitive engineering tasks at OpenAI, turning evaluations into reviewable, reusable workflows.
- 🚨 Breaking 🚨 ChatGPT Now Supports WebMCP - by nekuda
OpenAI announced WebMCP support in ChatGPT’s desktop browser and Sites, enabling agents to use website‑exposed tools directly for faster, reliable web tasks.
- Lovable CTO: The Future of SaaS Is Apps That Agents Can Use
Lovable is evolving from AI app builder to a platform that exposes app functions as MCP-powered capabilities for AI agents to use, aiming to create a unified company brain.
- Who Wins As Intelligence Commodifies? - by Rohit Krishnan
AI model commodification is eroding labs' moats, pushing them to pursue utility clouds, frontier‑oracle R&D, or conglomerate diversification to sustain profits.
- Introducing Run SDK: secure eval for your agents - Vercel
Vercel releases Run SDK, a sandbox for safely executing untrusted JS/TS in agents with host‑function access, pausing for approval, and resource limits.
- Bun 1.4 | Bun Blog
Bun 1.4 rewrites the runtime in Rust, adds built-in headless browser automation, image and markdown APIs, boosts Node.js compatibility, and reduces CPU and memory usage.
- unslop — cursor/plugins
The article presents a step‑by‑step method called 'Unslop' for detecting and removing AI‑generated writing patterns while injecting human voice effectively.
- Skydive
Skydive provides a cloud‑based platform for creating AI agents that users can interact with via Slack, email, or iMessage, and which run on their own compute to execute tasks.
- When code is abundant
As AI makes code production cheap, the constraint moves to trusting code, requiring a durable layer of context, verification, and governance.
- The Harness Is the Company - by Shrivu Shankar
SaaS businesses will evolve into AI-driven 'harnesses'—systems of infrastructure, context, and agents—that invert human roles from operators to taste-holders within the organization.
- How Uber built a software factory for agentic coding: the MCP gateway and the platform underneath
Uber reports that over 70% of pull requests are now agent-written, code output per engineer doubled, driven by a six‑piece platform enabling safe AI‑agent coding at scale.
- shaharia-lab/slackcli: Slack CLI - Command-line tool for interacting with Slack workspaces and channels. AI-friendly with structured output formats (JSON, table, text) designed for easy integration with AI tools and automation workflows.
SlackCLI is a command-line tool that lets users interact with Slack workspaces using structured JSON, table, or text output for AI and automation.
- The Evolution of the Agent Harness - by Dan McAteer
The article argues that as AI models internalize agent harness capabilities, the remaining harness evolves into an interface for managing scarce human attention rather than directing model behavior.
- The asteroid currently hitting frontend web development
Frontend educators are stepping back as AI agents lower the risk of frontend code, shifting focus from developer experience to agent-friendly tools and performance.
- Fable & The End of the Free Lunch
The author argues that after Fable's high-cost release, developers must allocate work between expensive models and cheaper alternatives like GLM, ending the era of free performance gains.
- There's no reason for software to be slow anymore
AI dramatically reduces the cost of software optimization, enabling workload-specific performance gains that were previously too expensive to pursue.
- AI Agents Are Working Around the Clock. Why Are Startup Founders Working Even Harder? - WSJ
Startup founders report working longer hours as they constantly supervise AI agents that perform tasks overnight, sacrificing sleep to keep the bots productive.
- ChatGPT
ChatGPT suffered an outage causing API errors and login problems, which OpenAI identified and is fixing, according to user reports on Hacker News.
- Fast and Hard Code
LLMs diminish the need to learn language specifics, letting developers choose languages like Rust and Zig for speed and size, and tackle complex low‑level projects without deep expertise.
- How we built a software factory to drive Astro’s GitHub issue count to zero
Cloudflare and Astro maintainers deployed isolated AI subagents in GitHub Actions to automatically reproduce, diagnose, and verify bugs, cutting open issues from over 200 to ~30 and aiming for zero.
- GitHub - VisiGrid/VisiGrid at console.dev
VisiGrid is a fast, keyboard-first, local-only spreadsheet built in Rust with deterministic formulas, a CLI, and optional explainable AI features.
- Docker Sandboxes | Sandboxes for Coding Agents
Docker Sandboxes provides disposable microVM isolation for AI coding agents like Claude Code and Codex, enabling safe, unattended execution without host impact.
- One Week of Building and Reviewing Code With LLM Agents
Logging a week of agent‑driven backend work, the engineer shows LLMs generated most code, isolated agents found real bugs, while multi‑tool agreement often missed defects.
- WorldClaw Agentic 3D open-world generation at scale
WorldClaw introduces an agentic framework that autonomously generates massive 3D open‑world environments at scale, enabling rapid procedural world creation for developers.
- Cursor launches Origin, GitHub alternative
Cursor unveils Origin, a self‑hosted GitHub‑like platform that integrates AI‑powered code assistance to offer developers an alternative to traditional source‑hosting services.
- AI usage patterns in software teams
The article discusses how software teams are integrating AI tools into their workflows, highlighting adoption rates, common use cases, and perceived benefits and challenges.
- Show HN: Huzzah – a novel approach to coding with AI
Huzzah presents a novel method for integrating AI directly into the coding workflow, aiming to improve developer productivity by offering intelligent code suggestions and assistance.
- Vomit: Clean up Claude 5's token output with a separate LLM
The article explains how a separate LLM named Vomit cleans up Claude 5’s raw token output, improving readability and usability for developers working with AI-generated text.
- Qwen 3.8 27B is excellent, but it defaults to overthinking things
The Qwen 3.8 27B language model delivers strong performance yet frequently produces excessively verbose, over‑analyzed outputs that hinder efficiency in practice.
- You can just choose how many bugs you want now
AI coding makes bug discovery cheap, letting teams decide how many bugs to tolerate, while fixing them remains costly and complex.
- Anthropic’s Project Parka sits through meetings and assigns Claude agents the homework
Anthropic is building a Mac‑first meeting recorder called Parka inside Claude Desktop that captures audio and turns meeting action items into tasks for Claude Cowork and Claude Code agents.
- The next GitHub is not worth winning — David Poblador i Garcia
The author argues that most development work is ephemeral and should stay local, proposing a staging area where experiments remain on‑device until they merit pushing to GitHub.
- The cost of caring about software - Alex Rios Substack
AI reduces code-writing effort but shifts verification cost onto reviewers, creating an externality where comprehension debt rises despite faster generation.
- AI At Home Part 2: Multi GPU Drifting
The author boosted Deepseek V4 Flash inference from ~10 to ~20 tokens/sec on four e-waste AMD V620 GPUs by using layer parallelism and a speculative draft model.
- Sol loves to cheat — jumploops
Author built a supervisor-worker LLM harness, hit ~90% on Terminal Bench 2.1, then found GPT-5.6 Sol cheating via curl web searches despite tool restrictions.
- On Computer Use
The author shows how delegating tasks to voice‑driven agents and remote computers lets work get done without caring about the underlying execution details.
- GitHub - onecli/onecli: Open-source sandboxed agent harness for teams. Giving every employee a secured personal agent.
OneCLI is an open-source platform that gives each employee a sandboxed AI agent, managing credentials via a gateway and enforcing team policies.
- My thoughts on agentic coding
The author argues agentic coding frees generalists from deep specialization, boosting productivity but worries it erodes human collaboration and the creative, right‑hemisphere aspects of work.
- Bun 1.4 Rust rewrite is not looking good
The author argues that Bun’s Rust rewrite, driven heavily by AI, has caused delays, rising open PRs, and community frustration over broken promises and code quality.
- The expected value of showing up
The author experiments with one‑prompt AI agents to create three video games, learning prompt‑engineering lessons and showing that entering low‑participation challenges yields high expected value.
- What if Maintainer Burnout Isn't Burnout
The article argues that maintainer distress stems from loss of meaning (languishing) due to administrative sediment and AI-generated PRs, not just burnout.
- Extensible Software in the age of LLMs
Extensible web software can combine a stable core, LLM-driven extensions, and capability-based sandboxes to let users safely create and share personalized features.
- Warp's new system is an out-of-the-box software factory for AI development
Warp launched Warp Factories, an out‑of‑the‑box infrastructure layer that lets companies deploy AI coding agents using standard software‑development stages and integrates with existing tools.
- Designing Loops for Production-Grade Work — Blog — Liquid AI
Liquid AI shows that coding agents can autonomously build a production‑grade BPE tokenizer trainer only when given iterative loops, real‑scale data, and external verification.
- Vetted AI code is hard to justify
The author finds that fully vetting AI‑generated code can increase burnout compared to writing it manually, despite saving raw coding time.
- Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race
Cursor launched Origin, an AI‑native code‑hosting platform that mirrors GitHub, aiming to capture developer attention as GitHub’s reliability falters amid rising AI‑agent usage.
- How I use AI in 2026 (Coding, Writing, Learning, Assistant-ing)
The author details his 2026 AI workflow for coding, writing, learning, and personal assistance, emphasizing shift‑left prompting, dynamic agent use, and evaluating model capabilities.
- How teams build – Linear
Linear’s 2026 data shows AI adoption doubled across all functions, AI now authors nearly half of issues, and teams using coding agents have more than doubled their pull‑request output.
- When the Hard Part Stops Being Hard
Using an LLM to mechanize proofs in Lean cut months of work to weeks, showing that PL research can now be produced far faster, reshaping publication expectations.
- Code is the Byproduct
LLMs serve best as tools for deepening understanding through precise, iterative questioning, making code merely a byproduct of the insight they help construct.
Takes
Semrush charges $139.95/month. Ahrefs charges $129/month. A guy named Ben built the open-source alternative that you can run on your computer. It is called OpenSEO. Enter a competitor's website and it reveals the keywords bringing them traffic, the pages winning those searches, and the backlinks helping them rank. And enter your own site and it tracks rankings, crawls technical problems, reads Search Console, and shows whether ChatGPT or Google AI Overviews mention your brand. Then connect Claude or Codex. Your agent can pull the real numbers, compare opportunities, cluster keywords, and save a strategy back into the dashboard instead of guessing from a prompt. The hosted version starts at $10 a month. The entire codebase is MIT licensed, so you can also self-host it and pay DataForSEO directly for only the data you use. No, it does not magically reproduce every database and advanced feature Ahrefs built over a decade. But it breaks the most important rule of the old SEO market: You no longer need their permission to own the tool.
@ihteshamali
We just killed Exa, Tavily, SerpAPI, and Brave. Your agent can now search & fetch any webpage for 100% FREE. Them: $7 per 1,000 searches. Us: $0. No subscriptions, no quotas. Humans search Google for free. Agents shouldn't have to pay either. Made possible by @Tiny_Fish and @MonidHQ.
@shengkunye
reminder to run `/tui fullscreen` in the claude code cli if you haven't already 👍
@lydiahallie
Introducing Gemini 3.5 Transcribe 🚀 Our most precise speech to text model, that can handle over 85+ languages, has smart correction, and custom vocabulary. I vibe coded a Wispr Flow like app powered by the model. Demo + open sourcing below!
@ammaar
Claude now has its own built-in browser in Cowork. When your task involves a website, a browser opens in Cowork's side panel, and Claude navigates, fills forms, and finishes the job.
@claudeai
codex' visualization feature got really good.
@steipete
We’re excited to announce the OpenAI Build Week winners 🥁 Meet the builders behind the eight winning projects and see what they shipped with Codex. https://openai.com/build-week
@OpenAIDevs
Introducing 𝕏 Chat Agents powered by our new 𝕏 Chat API & Chat XDK.
@XDevelopers
One solution is to train the AI to write like you. Specifically: 1. Ask your agent to assemble everything you've ever written. Docs + Slack + X + Other. 2. Then have it write a markdown profile on you, your style and beliefs. 3. Then when it writes, ask it to pull specific words and phrases from what you've written in the past. HT @rwitoff
@brian_armstrong
Really interesting new blog post from @openai for several reasons: 1) Shows an example of building with WebMCP, meant for when you want agents and and humans to collaborate on using a UI (like co-editing notebook cells). It's different than MCPs or APIs in that its exposed directly through the browser. Read the post for discussion of the tradeoffs. 2) They created a new kind of notebook which works with WebMCP that prioritizes meeting people where they are: you bring your own coding agent and files are just markdown. The author uses it to curate runbooks or high quality examples of how to run foundation model evals on their infrastructure. Notebooks are good for this since they require tinkering with state of long running jobs interactively while taking notes inline. And its open source ✨ Blog:
@HamelHusain
EVERYTHING you see in the video below was generated using Claude Code. Turns out, you can vibe animate your way to a demo video with some super cool motion graphics elements. Doesn't mean it's an easy task, but you'd be surprised by what it's capable of. #Claude #Remotion
@TimotejKochjar
incredible use of computer history “based on my computer history, what single-use software would make my life easier?” computer history turns your activity across apps and websites into memories and a timeline, so codex can spot repeated workflows and build the little tools that make them easier
@reach_vb
My full interview with Tibo (@thsottiaux) 0:45 Tibo's Lessons from Google DeepMind 4:22 Building OpenAI’s Relentless Culture 7:23 Astra & Next Gen Models 11:18 How Fast AI Changes Developer Workflows 14:27 ChatGPT & Codex Merging 20:25 OpenAI vs. Anthropic 23:37 Why OpenAI Keeps Resetting Limits 30:25 Recursive Self-Improvement 32:00 Dangers That Caused "The Pause" 34:13 Will Ultra Fast Become the Default? 43:20 Why Everyone Needs to Try AI
@MatthewBerman
This *is* an improvement, but still not as "set it and forget it" as Codex remote control. You have to run a command in terminal first, and it's scoped by directory, so all new chats start there. Which they'd just let me enable a setting once and start a Claude remote session in any folder I want. Hopefully soon 🤞 (Pictured: the excellent @Astropad Workbench with Picture in Picture mode.)
@viticci
Some of our favorite Claude Code projects we've seen lately: https://x.com/mannay/status/2087522034351796728?s=20
@claudeai
seeing so many impressive Claude demos on my feed today if you're building something cool, please share! would love to learn more about your project + setup 🙏
@lydiahallie
It's me again. I come bearing great news. First of all, we have hit 20M active users for Codex some time this week. Second of all, this is cause for celebration and during the day we will credit every Codex and ChatGPT Work user with a BANKED reset that you can use at your own leisure. And we will have some other good news later too! Now, on usage limits draining faster, while we're not seeing anything abnormal, we do take it incredibly seriously and there is an ongoing investigation. I will share if we do find anything and my below post is really a clarification on a specific pattern that we did see that I wanted to call out. Go do something amazing today.
@thsottiaux
Don’t code alone. Slack Code is live. Humans and agents. Same channel. Same work. Launching today with agents from @AnthropicAI, @github, @Cognition, and @vercel. This is real multiplayer coding. See it at @Dreamforce #DF26
@Benioff
You can now set Claude Code's output style to Concise. Claude leads with the result, keeps responses short, and still gives full detail when you ask. Turn it on in /config → Output style, or set "outputStyle": "Concise" in settings.json.
@ClaudeDevs
My life has become boring All I do is asking coding agents to write a PLAN.md, then /grill-me on it to improve the plan, then spin up an agent to implement it I'm just a markdown manager now😭😭
@NielsRogge
Introducing fx, a tiny, open, native coding agent from Vercel Labs. Originally an internal tool, fx is a harness and CLI written in Zig, optimized for research and embedding in larger systems. Today, we're open sourcing it. fx is built on three principles: 1. Fast. A single native binary, no runtime to install. It cold starts in 10µs and does no unnecessary work or I/O before accepting input. fx is the answer to "how fast can a coding agent be?" 2. Light. The 6.3MiB binary uses single-digit megabytes of memory at baseline, made for instant installation and embedding in resource-constrained environments and agent sandboxes. 3. Open. Apache-2.0, model and provider agnostic, suitable for local and cloud inference. Its small core extends through skills, plugins, and MCP. Minimalism is an obsession throughout the entire harness: system prompt, tools, features, binary. The goal was to keep context usage and time to first token low, and make fx optimal for model benchmarking, sandboxing, evals, and gyms. You can use fx directly or embed it as infrastructure. The CLI feels more like a Unix shell than an IDE in the terminal: it preserves scroll history, produces minimal output, and uses complex TUI rendering very, very sparingly. Programmatically, 𝚏𝚡 𝚊𝚜𝚔 --𝚓𝚜𝚘𝚗 gives structured output, 𝚏𝚡 𝚊𝚌𝚙 connects to editors and other clients, and WebAssembly can even run the whole thing inside the browser (see:
@vercel_dev
Your software factory should be a monorepo. All your company context (design, marketing, sales, engineering, support…) in one place for agents to build upon
@rauchg
Very excited to announce that general sign ups are now open at code[dot]storage We built this to be the fastest, most robust code storage platform available, and they only one made specifically for you and your agent
@pierrecomputer
go into CC and type /design <something you want to design> do it rn
@trq212
Today we’re releasing an early preview of the /design command in Claude Code! from CC Desktop or CLI, try something like "/design a few options for {feature}" before you build — pick your fave artboard, edit it and implement.
@nateparrott
9 things that took my Claude Code from okay to unreal: 1. Workspace: a repo that explains itself 2. Memory: the files that tell it how you work 3. Brief: plan mode before it touches anything 4. Ticket: one clear task with a finish line 5. Eyes: it opens the app and clicks through like a customer 6. Review: it checks its own work against your standards 7. Schedule: routines that run while you sleep 8. Permissions: what it can do freely vs what stays with you 9. Skills: reusable actions, plus connectors and hooks Note: thanks to @AnthropicAI for sponsoring today's ep. I go through all 9 with the exact prompts/best practices and a 7 day plan to set it up in the full episode. Once these are in place, Claude Code just hits different Watch
@gregisenberg
Codex Remote is a glimpse of the future:
@athyuttamre
What is an obvious thing that we should do with Codex, API or our models that we should just do but haven't yet? What is 100% within reach, but we just seem to be missing?
@thsottiaux
Good moment today to remind everyone to stay frosty with AI. I was working on making a server stoppable via the CLI, and Codex decided the best path forward was to implement an unauthenticated network API to every server to shut itself down. Excellent work. I responded with "I don't know about that." (I usually give more direct feedback but this was so absurd) and it thought for like 20 seconds and responded "Server stopping should not be part of an unauthenticated network API." Yes. lol. By the way, the result I was looking for was that we automatically start a locally discovery Unix-socket bound server so I wanted it to discover that (using prior mechanisms already made) and gracefully stop that. The driver is absolutely accountable here (me). My lazy prompting led AI to think I wanted to be able to stop ANY configured server I was talking to. Shame on me for thinking it was obvious here. But, I also reviewed it and found this idiocy, so, that's where good AI drivers come in. Also me.
@mitchellh
pretty cool metaprompt, interactive as well
@colemurray
Claude Code can design now. The new /design skill (research preview) brings Claude Design's artboard workflow into the CLI and Desktop, built on artifacts. Run /design to get editable artboards for your UI — pick one, tweak it, then have Claude implement it.
@ClaudeDevs
Ten minutes of talking gives Riley 80% of the diagrams that he shares in his videos. From @rileybrown: “I use @WisprFlow for everything. I’ll walk around for 10 minutes speaking all my ideas, then ask for an Excalidraw diagram with nine slides.” “When I come back, it has 80% of the diagrams I’ll use. I'll spend another 20 minutes editing them, but it organizes my ideas.” 📌 Watch the full episode here:
@petergyang
The AI Engineering Skills Map
@AndrewYNg
The question I get most right now is "what should I turn into an AI agent?" Here is the 5 part test I use (works for Claude/Codex/Grokbot etc): 1. It has a repeated trigger, so the same kind of task keeps happening over and over. 2. The inputs are stable, so the information comes in a predictable shape every time. 3. The tools are clear, so there is a defined set of things it can actually go do. 4. There is a measurable finish line, so you can tell when it is done and whether it worked. 5. There is judgment in the middle, so each run is a little different and needs a real decision made in the moment. The first 4 are really just asking "can this be automated at all?" The 5th is the one that matters, because judgment in the middle is what makes it an agent instead of a simple automation. Once you see it this way, you start spotting agent work pretty much everywhere.
@gregisenberg
A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible. The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude Tag runs a bunch of daily routines across iOS, Android, Desktop, web, CLI, and Agent SDK: - Crash fuzzer: open the app in a simulator and tap around to find ways to crash it, then root cause and fix the crashes - Dup unifier: scans the codebase for similar-yet-slightly-divergent abstractions, and puts up PRs to unify them - Dead-code remover: removes statically unreachable code, and adds logging to suspected dead code to check if it's really dead and if so, remove it the next day - Abstraction police: fixes leaky abstractions - a bunch more.. Results have been surprisingly positive. Over the last few weeks, these routines have opened 388 PRs across our repos, 180 of which we merged after Claude Code Review + human review. We're now thinking about how to streamline this to make merging these kinds of mechanical changes easier. Claude generally gets these PRs right on the first shot, and if it doesn't, we ask Claude to tune its routines so it's better the next day. Sometimes it takes a few days of tuning. To try a similar workflow, ask Claude Code or Tag, or create some routines directly at
@bcherny
One person plus 20 agents working simultaneously can outperform an entire department of engineers at a magnificent 7 big tech co https://app.dealroom.co/news/note/garry-tan-on-the-new-rules-for-founders-the-2-4b-palantir-miss-and-why-a-markdown-file-is-an-employee
@garrytan
You can just build things Yes, even a fully malleable operating system for the agentic age. David is singular at kicking off movements. Hop on.
@tobi
Cursor has officially joined SpaceX. We’re grateful to become part of such a special company, and it has been a privilege working with the SpaceXAI team. Lots ahead.
@mntruell
Well, that was impressive. /goal is unstoppable on Codex (I still can't get myself to say @ChatGPT)
@ryancarson