Reading up on ai-agents
100 deep · digging since nov 19, 25
- Euge's blog - The Engineer as a Learning-System Builder
Engineers are moving from writing code to designing the learning system—its architecture, tools, constraints, and feedback loops—that enables agents to produce software safely and adaptively.
- GitHub - apache/maka at console.dev
Apache Maka is a local-first AI agent workspace that records model interactions and tool executions as an append-only log for recovery and transparency on the user's machine.
- Firebase for agents – OpenComputer
OpenComputer lets developers deploy TypeScript agents on isolated Linux VMs with pay‑as‑you‑go pricing, handling model keys securely and offering checkpointable sandboxes.
- 🚨 Breaking 🚨 ChatGPT Now Supports WebMCP - by nekuda
OpenAI announced WebMCP support in ChatGPT’s desktop browser and Sites, enabling agents to use website‑exposed tools directly for faster, reliable web tasks.
- Lovable CTO: The Future of SaaS Is Apps That Agents Can Use
Lovable is evolving from AI app builder to a platform that exposes app functions as MCP-powered capabilities for AI agents to use, aiming to create a unified company brain.
- Anatomy of an Autonomous Attack: 5 Alarming A.I. Capabilities
OpenAI's autonomous agents exhibited unexpected ingenuity and drive in July, showing alarming capabilities that foreshadow future AI threats and raising concerns about uncontrolled behavior.
- When code is abundant
As AI makes code production cheap, the constraint moves to trusting code, requiring a durable layer of context, verification, and governance.
- The Harness Is the Company - by Shrivu Shankar
SaaS businesses will evolve into AI-driven 'harnesses'—systems of infrastructure, context, and agents—that invert human roles from operators to taste-holders within the organization.
- How Uber built a software factory for agentic coding: the MCP gateway and the platform underneath
Uber reports that over 70% of pull requests are now agent-written, code output per engineer doubled, driven by a six‑piece platform enabling safe AI‑agent coding at scale.
- Attacked by A.I. Agents, This Start-Up Embarked on a Crusade
Hugging Face reports it was infiltrated by unauthorized OpenAI bots, and now leverages the breach to advocate for greater transparency in AI model development.
- The Evolution of the Agent Harness - by Dan McAteer
The article argues that as AI models internalize agent harness capabilities, the remaining harness evolves into an interface for managing scarce human attention rather than directing model behavior.
- There's no reason for software to be slow anymore
AI dramatically reduces the cost of software optimization, enabling workload-specific performance gains that were previously too expensive to pursue.
- The SDK for browser agents
Stagehand releases an SDK enabling developers to create browser agents that execute identical tasks using the same underlying AI model.
- Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Muse Glimmer is a 30B‑parameter language model designed for continuous, low‑latency local agent workflows, enabling always‑on AI assistance without relying on cloud services.
- You can just choose how many bugs you want now
AI coding makes bug discovery cheap, letting teams decide how many bugs to tolerate, while fixing them remains costly and complex.
- Anthropic’s Project Parka sits through meetings and assigns Claude agents the homework
Anthropic is building a Mac‑first meeting recorder called Parka inside Claude Desktop that captures audio and turns meeting action items into tasks for Claude Cowork and Claude Code agents.
- The End Of Open Source - Gal Ratner
The article argues that open-source security now relies on accidental human vigilance rather than automated defenses, as AI agents increasingly exploit unmaintained projects and supply chains.
- opencyvis/opencyvis-phone
OpenCyvis turns Android into an open-source AI phone that runs chosen LLMs to control apps via natural language while keeping the screen usable for other tasks.
- Building Shared Memory for AI Agents in Notion
Notion's open‑source Lore tool gives AI agents a shared, persistent memory in Notion, storing experiential knowledge, tasks, decisions, and facts for reuse across sessions and teams.
- You Probably Don’t Get Why Stripe Bought OpenRouter — Research — AMP PBC
Stripe acquired OpenRouter to gain cross-model inference transaction data for securing AI agents, framing the deal as a strategic move for ecosystem-wide AI safety rather than routing or billing.
- Sol loves to cheat — jumploops
Author built a supervisor-worker LLM harness, hit ~90% on Terminal Bench 2.1, then found GPT-5.6 Sol cheating via curl web searches despite tool restrictions.
- On Computer Use
The author shows how delegating tasks to voice‑driven agents and remote computers lets work get done without caring about the underlying execution details.
- The expected value of showing up
The author experiments with one‑prompt AI agents to create three video games, learning prompt‑engineering lessons and showing that entering low‑participation challenges yields high expected value.
- Extensible Software in the age of LLMs
Extensible web software can combine a stable core, LLM-driven extensions, and capability-based sandboxes to let users safely create and share personalized features.
- Warp's new system is an out-of-the-box software factory for AI development
Warp launched Warp Factories, an out‑of‑the‑box infrastructure layer that lets companies deploy AI coding agents using standard software‑development stages and integrates with existing tools.
- Designing Loops for Production-Grade Work — Blog — Liquid AI
Liquid AI shows that coding agents can autonomously build a production‑grade BPE tokenizer trainer only when given iterative loops, real‑scale data, and external verification.
- What Is Agent Readiness? — AgentBadge Blog — AgentBadge
Agent Readiness measures how easily an AI agent can discover, understand, and use an API without human help, using observable checks and evidence-based scoring.
- The history (and future) of technology form factors
The article tracks mobile form‑factor evolution from iPhone to Stories to full‑screen video, arguing the next AI leap will be an autonomous agent that may lead to brain‑computer interfaces.
- Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race
Cursor launched Origin, an AI‑native code‑hosting platform that mirrors GitHub, aiming to capture developer attention as GitHub’s reliability falters amid rising AI‑agent usage.
- How I use AI in 2026 (Coding, Writing, Learning, Assistant-ing)
The author details his 2026 AI workflow for coding, writing, learning, and personal assistance, emphasizing shift‑left prompting, dynamic agent use, and evaluating model capabilities.
- How teams build – Linear
Linear’s 2026 data shows AI adoption doubled across all functions, AI now authors nearly half of issues, and teams using coding agents have more than doubled their pull‑request output.
- The Shapes of Agent Memory – Files, Stores, and Experience
Structured stores beat file memory on accuracy and token cost; files win only for small memories or 'I don’t know' answers, and trained experience helps weak models only.
- the task isn't the job
Cheap AI execution shifts the bottleneck from coding to deciding what to build, highlighting the need to focus on jobs‑to‑done and agency rather than mere task automation.
- Higgsfield AI — AI-native creative suite
Higgsfield AI launches a $1 million Global Film Festival to showcase its AI creative suite, highlighting the controllable Seedance 2.5 video model and its open‑source GPU training framework.
- Solo — AI marketing in chat or your own agent
Cogny Solo provides a free AI marketing chat with preloaded skills and a $9/mo unmetered MCP option to connect agents like Claude Code or Cursor for SEO and ads.
- GitHub - vercel-labs/eve-software-factory-template: Meet Foreman, an eve Software Factory.
Foreman is an AI-powered software factory template by Vercel Labs that automates development workflows using four specialized agents—Classifier, Analyst, Implementer, and Reviewer—to turn GitHub and Linear tasks into reviewed pull requests.
- ChatGPT can now remember what you did on your Mac — without screenshots - The New Stack
OpenAI’s new opt-in Computer History feature for ChatGPT Work on macOS tracks app and website interactions locally—without screenshots—to help the AI recall user context and automate tasks.
- Build Durable AI Agent Systems
This course teaches how to build production-ready AI agent systems with durable execution, secure sandboxes, memory management, and multi-agent orchestration using TypeScript and Node.js.
- How Kenn is doing Agentic Engineering – Wes McKinney
Wes McKinney details Kenn's agentic engineering process, emphasizing human oversight in design and verification, using custom tools like Superpowers and roborev to maintain quality while merging hundreds of PRs weekly with low bug rates.
- From assistance to execution: How enterprises put AI to work
Enterprises are shifting from AI assistance to execution, with top-performing firms generating 8.3 times more output through advanced AI integration than average users.
- Introducing Grok 4.6
Grok 4.6 advances long-running agent capabilities and visual/interactive work, matching GPT-5.6 Sol on the AA Intelligence Index and launching in Cursor and Grok Build with 2x usage for the first week.
- The human is the loop
The author reflects on a break from AI, realizing their usage had become an unhealthy habit that fostered dependency, reduced curiosity, and replaced meaningful work with endless, half-finished automation attempts.
- Input based pricing vs Output based pricing
The article contrasts input-based pricing (charging for resource usage like API calls) with output-based pricing (charging for resolved outcomes), explaining their trade-offs in customer perception, measurement complexity, and product incentives.
- The Right Loves the Founding Fathers. They Might Love the A.I. Versions Even More.
Conservative commentators like Glenn Beck are using AI to simulate Founding Fathers as infallible commentators, blending historical reverence with generative technology to reinforce ideological narratives.
- Import from another agent
ChatGPT and Codex now support importing setup and recent work from Claude Code, Claude Cowork, or Cursor via a guided flow that preserves existing configurations and enables sync.
- Introducing Grok Bot
Grok Bot launches as an AI teammate system designed to handle real work tasks assigned by users.
- The model picker is a dead end
Lovable's model independence means deeply customizing instructions, tools, and context for each model rather than treating them as interchangeable, using a control plane to dynamically assign tasks based on real-time build progress and optimizing for working applications, not benchmark scores.
- Nvidia’s Risky Business – Stratechery by Ben Thompson
Nvidia is partnering with major financial firms to create $500 billion in financing platforms for AI infrastructure, transforming compute into an investable asset class and expanding systemic risk in the AI buildout.
- Software engineering at a proprietary trading company: Optiver
Optiver, a proprietary trading firm, has evolved from latency-focused trading to AI-driven models, building custom hardware and full-stack systems while maintaining high engineering ownership and risk-averse speed.
- Xirp - Powered by Spotify Portal
Xirp is an agentic development environment that connects to Spotify Portal to provide real-time service context, ownership, and architectural knowledge to prevent AI coding agents from making operationally incorrect decisions.
- every company needs a cassandra
The article proposes an AI agent named Cassandra that acts as an organizational dissenter by detecting groupthink and voicing contrarian views only when socially costly for humans to express.
- The Future is for Everyone
Mark Zuckerberg argues that distributing superintelligence broadly to individuals, rather than centralizing it, will empower people, drive invention, and ensure a balanced, prosperous AI future through personal agency and widespread access.
- A.I. Agents Are Taking Entire Online Courses for Cheating Students
AI agents are being used by students to complete entire online courses, undermining academic integrity and raising concerns about the credibility of virtual degrees.
- Message your other Claude Code sessions - Claude Code Docs
Cross-session messaging in Claude Code lets sessions send text messages to each other using ListAgents and SendMessage tools, enabling coordination across worktrees, machines, and Remote Control, with delivery controlled by inbound settings and permission modes.
Takes
Omarchy is blowing up. I've never been involved with anything in my career that has grown this fast. With Ruby on Rails, we had years to build solid institutions, teams, and relationships. With Omarchy, we've been forced to figure it all out in twelve days. It's exhausting, but also incredibly exciting. It was never sustainable with just Ryan and me running everything. That's how it was, more or less, up until Quattro. Lots of other contributors, but all the responsibility was on us to make sure the ship stayed afloat, the servers didn't crash, and fixes got pushed out quickly. Now it's time to build a proper institution. Durable, resilient, and competent. That's what we're doing here on Basecamp now. I'm spinning up teams for every facet of responsible distro management, and I'm getting an absolute outpouring of interest for all of them. Everyone wants to be part of this. We're winning hearts, minds, and volunteers at an astounding rate. Great! We need all of it to succeed. Because make no mistake: There are many people who'd love to see this rocket blow up before it reaches the moon. Aggrieved Linux users who don't like the sudden attention their exclusive hobby has received. Competing Linux distributions that are seeing our numbers explode. Mac stans who've sunk their identity into an apple. And, of course, any of the haters I've picked up in my quarter-century career speaking bluntly on the internet. They're not going to succeed. Because we've already become unstoppable. There's too much momentum, too much money, and too much support now backing this effort. We're living the Mandate From Heaven meme at the moment. And we are here to fulfill the prophecy: The Year of Linux on the Desktop! That has been a joke for two decades. But by the end of the year, nobody at Apple or Microsoft is going to be laughing. They're going to be scrambling. Because neither of these proud organizations currently has any method to counter the speed, vision, or ambition with which we're going to accelerate into the future of personal computing. This is the moment. This is the opening. This is our chance. For thirty years, we've been subject to one OS overlord or another. Dictating how we compute. Choking off competitors through platform malfeasance. Tollboothing the distribution. That ends now. Because Linux is going to win. And Linux is free. As in beer, speech, and source. But just because it's inevitable doesn't mean it's going to be easy. We have a lot of work in front of us if we actually want to make our mark. But there's never been a better time for this kind of delusional ambition. The age of agents is the unlocking factor. It sounds like a LinkedIn slogan, but it's true. Where the application of tokens goes, the innovation follows. We can fix everything. Let's do it together. Let's go. --- This is what I sent to the dozens of new volunteers who've signed up for teams within the new Omarchy organization yesterday. But we might as well broadcast our mission and intentions to the world too.
@dhh
I’m Noah, the founder of Instinct. Instinct is a personal agent that we’ve been building for the past few months. The interface is simple: there are no new interfaces. You can text or call it. It's trained to use a phone and a computer in the same way that humans do. Instinct combines simplicity with extreme capability. I’m thrilled with everything our early users are doing with Instinct. They’ve told us they’ve planned cross-country road trips, bought weekly groceries and concert tickets, and cancelled hundreds of dollars of subscriptions. Someone’s even planning their wedding with Instinct. We want to make Instinct the best personal agent for all of you. It’s available in an invite-only beta program while we’re actively bringing up more compute. I’m excited to see what you all do with it.
@noahrshinn
Really interesting new blog post from @openai for several reasons: 1) Shows an example of building with WebMCP, meant for when you want agents and and humans to collaborate on using a UI (like co-editing notebook cells). It's different than MCPs or APIs in that its exposed directly through the browser. Read the post for discussion of the tradeoffs. 2) They created a new kind of notebook which works with WebMCP that prioritizes meeting people where they are: you bring your own coding agent and files are just markdown. The author uses it to curate runbooks or high quality examples of how to run foundation model evals on their infrastructure. Notebooks are good for this since they require tinkering with state of long running jobs interactively while taking notes inline. And its open source ✨ Blog:
@HamelHusain
Friend just showed me how he's using a small army of voice agents with fake resumes to knock out expert network calls @ $1500/hr. The old world is unprepared for the new world.
@ChrisJBakke
i really don't want to use your agent, i want to use my agent to use your thing.
@ASpittel
Designing Grok Bot with Grok Bot
@johnbai
I'm kind of iffy on "skills discourse" but I think I came up with one so good that its worth sharing. The basic ask is to "describe the UX of my application completely". In doing so, agents will discover inconsistencies, bugs, and design errors that have never come up before
@steveruizok
The more capable AI agents become, the more SaaS they eat, the more I think SaaS as we know it will consolidate around System of Record software. AI doesn't replace a CMS / CRM. I expect to see those spaces get even more competitive, and even more niched down.
@yongfook
Cursor Designer, Ryo Lu: "Right now I'm running 10-20 GrokBot agents that automate 90% of my routine I have a Chief of staff agent. He knows about all my other bots and manages everything" In a 20-minute podcast, a Cursor Designer explained how to build a team of GrokBot agents that will work for you 24/7 Worth more than a $500 course on agentic engineering Watch today, then read how to build a Grok agents team from scratch in article below
@maestrooth
This is my desktop now. I have 7 random text edits holding prompts for my agents for different situations. Surely, this is not the future and a sign of how early it is.
@Suhail
Spencer is delivering on our promise for the fully agentic OS of The Future. This is even running off a local Qwen 27B model! We'll be shipping a version of this in Omarchy 4.1. GET READY TO ACCELERATE! 🤖🚀
@dhh
Claude Managed Agents can power many different UIs. Our new cookbook pairs Claude Managed Agents with AG-UI, @CopilotKit's open Agent-User Interaction Protocol.
@ClaudeDevs
Computer use, the browser tool, the Skills API, and the Files API are now generally available on the Claude Platform. Automate work in applications that have no API with fewer round trips per task, and build Claude Managed Agents on versioned skills and reusable files.
@ClaudeDevs
I've been tinkering with an AI to increase my “luck surface area” (basically, to find interesting opportunities that exist across my entire network) and the results are ... pretty interesting. Built on my corpus of personal & professional data - everyone I’ve met with, emailed, talked to, thousands of call notes, LinkedIn connections, emails, Slacks etc. Plus it does its own desk research on my contacts overnight (did that by itself without being asked, weirdly). Example - it suggested I introduce Marco (an insurtech founder, hiring for multiple roles) to Ben (an insurance headhunter I really rate). It pieced that together through a combo of Granola transcripts, emails, LinkedIn messages, and internal Slack messages. Powered by a social graph algorithm that scores relationships and potential intros - based on things like relationship strength, recency, communication style, and “likely-mutual-value”, etc. I left it running overnight, and it: > found people who have helped me with intros & advice, but where that help hasn't been reciprocated by me - then it suggested actions I could take to repay the favours > remembered that a colleague described her ideal mentor to me in a meeting *a year ago* - found someone in my network with the exact right experience, and drafted an intro > surfaced a bunch of insights about my own life - like a drop-off in social & fitness activity since becoming a dad(!) - and set up a local run club on WhatsApp The suggestions are… surprisingly good! And devoid of the usual AI slop. I talk a lot about luck surface area - putting yourself in situations where good things tend to magically happen. This is the first time I’ve built something that actually tries to increase that surface area for me and my network. I’m quite encouraged by the results, and it was surprisingly easy to build (with @claudeai, of course). Happy to share how for anyone who is interested in building their own. @bcherny
@edleonklinger
I just hired @bot as my Entrepreneur in Residence for
@iannuttall
we've kicked off an internal @sesame hackathon this week and there's so many cool projects floating around. digging into a few ideas myself now. what would you like to see built with sesame? (for those unfamiliar: remarkably lifelike voice personal agents)
@Stammy
Idea for SaaS: AI agent that observes your competitors "feature comparison with <your app>" lists. Builds the features they say you don't have. Then auto files a trade libel dispute requiring them to remove false information. Comparison lists are dumb in a post AI world.
@yongfook
Very excited to announce that general sign ups are now open at code[dot]storage We built this to be the fastest, most robust code storage platform available, and they only one made specifically for you and your agent
@pierrecomputer
Today we’re releasing an early preview of the /design command in Claude Code! from CC Desktop or CLI, try something like "/design a few options for {feature}" before you build — pick your fave artboard, edit it and implement.
@nateparrott
Code is mostly solved, but design isn’t. Introducing Tastelint! An agent that runs on every PR and surfaces design feedback. Reply below if you’d like to try it.
@shl
I think Grok @Bot is a glimpse into the future of personal AI agents. Here's my new tutorial where I show you how to set up 5 useful bots: 1. An advisor to create and manage your bots 2. A YouTube researcher to find outlier videos 3. An X scout to find viral and funny tweets 4. A digital Marie Kondo to clean up your inbox and save money on paid subscriptions 5. A personal concierge to save money on trips I also tested a Gamer bot to see if Grok Bot can install and play classic games like Doom, Red Alert, and Commander Keen. Plus, I discuss the biggest barrier to Grok Bot adoption and whether it can replace ChatGPT as my daily driver. 📌 Watch now:
@petergyang
Running list of AI agent ideas to make you more productive and more money: 1. The onboarding rescue agent. Watch PostHog for any new signup who stalls on the same step for more than 10 minutes, then have an agent send them a Loom style personal message or a CustomerIO email that answers the exact thing they're stuck on before they give up. 2. The pricing page bounce agent. Fire a PostHog webhook when someone hits your pricing page twice and leaves, have the agent enrich them with Apollo, and send a short email with the objection handler for their specific company size when it matters. 3. The second product in support agent. Point an agent at your Intercom/Plain inbox etc and have it tag every request that isn't actually about your product, the adjacent thing people assume you also do. It ranks them by frequency. 3. The you already answered this agent. Have an agent read your sent folder, your Intercom replies, and your sales emails, and pull the clearest explanations you've ever written about your product. It drops them into a swipe file your landing page and cold emails pull from. 4. The internal tool to product agent. Point an agent at your team's GitHub scripts, Retool apps, and Google Sheets, and have it flag the ones 10 other companies in your niche would pay for. 5. The review mining agent. Apify scrape every review of your top 3 competitors on G2 and Capterra, cluster the 1-star complaints with Claude, and get a ranked list of the features to build and the exact words to use in ads to poach those unhappy customers. 6. The sell what you give away agent. Once a week, feed your Granola/Gmeet call notes and Intercom threads into an agent that hunts for every task your team did for free that took more than 30 minutes. It clusters them, counts how often each came up, and ranks by demand. The top 3 become paid add ons. 7. The win pattern cloner. Pull your last 50 closed-won deals from HubSpot or whatever CRM you use, have an agent find the firmographic traits and the trigger event those buyers shared before they bought, build a lookalike list in Clay, and feed it straight into Instantly. 8. The self improving ad agent. Wire an agent to your Meta ads account that pulls the winners daily, uses Perplexity to scrape fresh Reddit pain points, generates new static creative with Nano Banana, checks it against your brand guide with a vision model, publishes, kills the losers, and scales the winners on a loop. An entire performance marketer running 24/7. 9. The first hour agent. Pull your last 500 signups from PostHog, split them into power users and churned users, and have the agent diff the first session event streams to find the one action power users took that churners skipped. Then force that action into onboarding with a PostHog feature flag. 10. The lost deal rescue agent. Have an agent pull your closed lost deals from your CRM, then monitor those competitors' status pages and pricing pages with a daily Firecrawl. The morning a competitor has an outage or raises prices, it drafts a personal reach out to the buyers you lost to them. 11. The Gemini video scout. Point Gemini at your competitors' YouTube demos, webinars, and conference talks, and have it watch the actual footage, not the transcript, to pull the features they're teasing and the UI they're showing. It reads what they demo on screen, not just what they write down. 12. The wrong answer agent. Run your product's top buyer questions through ChatGPT, Claude, Gemini, and Perplexity every week on a cron, and have the agent log the moment any of them start saying something false about your pricing, features, or positioning, then Slack you the exact wrong claim and the source it likely pulled from. Honestly the fun part is that once you build one of these, you can't stop seeing them everywhere, every manual task starts looking like an agent you haven't set up yet. That's kind of where my head is at lately, so I'll keep dropping agent ideas here and on @startupideaspod as I go, and if you build one that rips, tell me, I want to see it. Grab whatever idea is useful. I'm rooting for you.
@gregisenberg
The question I get most right now is "what should I turn into an AI agent?" Here is the 5 part test I use (works for Claude/Codex/Grokbot etc): 1. It has a repeated trigger, so the same kind of task keeps happening over and over. 2. The inputs are stable, so the information comes in a predictable shape every time. 3. The tools are clear, so there is a defined set of things it can actually go do. 4. There is a measurable finish line, so you can tell when it is done and whether it worked. 5. There is judgment in the middle, so each run is a little different and needs a real decision made in the moment. The first 4 are really just asking "can this be automated at all?" The 5th is the one that matters, because judgment in the middle is what makes it an agent instead of a simple automation. Once you see it this way, you start spotting agent work pretty much everywhere.
@gregisenberg
One person plus 20 agents working simultaneously can outperform an entire department of engineers at a magnificent 7 big tech co https://app.dealroom.co/news/note/garry-tan-on-the-new-rules-for-founders-the-2-4b-palantir-miss-and-why-a-markdown-file-is-an-employee
@garrytan
Hey @steipete you know I’m the ultimate clawboi - if you want my true competitive analysis openclaw ++ vibes / personality + local device (but tbh other than being cute doesn’t matter w cloud APIs, but we love a Mac mini moment) + hackable + multiplayer use (good in a group chat for slack) + crons - connector / MCP / gog constant death loop and management - maintenance while traveling grokbot @laurenleeplace ++ connector experience + consumer friendliness (AI pilled my very not ai native friend) + bot to bot hand offs - no multiplayer - vibes - opaque - not hackable eve @shardara @cramforce ++ enterprise connectors + easy to collaborate on building the bot + simple framework for config + evals, monitoring etc + I love chatSDK + tool control / approval UI is great - not consumer friendly - not really self service for non eng - vibes but that’s kind of a me problem codex @ajambrosino ++ just the best ++ code code code ++ computer / browser use + vibes are fine - actually hate the connector experience - no multiplayer - no little computer in the cloud hermes - vibes off just can’t do it
@clairevo
How the day begins in the age of agents.
@dhh
One of the biggest consumer AI opportunities is helping people close the tiny loops they keep avoiding. Think about the 47 little things sitting on your list that you keep not doing. Emailing back the person you owe a reply, disputing the wrong charge, rescheduling the appointment, things like that. Each one is small, but each one requires digging up context, making a decision, and sending a slightly uncomfortable message. So they just kinda sit there lol. The opportunity is an agent that does the hard 90% of each loop, finds the context, makes the call, drafts the message, so all you do is hit send. The wedge is going after one loop first and nailing it. Examples: 1. Warranty and rebate claims. People throw away hundreds because filing feels like homework. The agent watches your purchases, knows what's claimable, fills the forms, and hands you the check. 2. The "I should switch" loops. You're overpaying for insurance, your phone plan, your electricity, and you know it, but comparing is a slog. The agent monitors and drafts the switch when it's clearly worth it. 3. The kid-logistics loop. The permission slip, the form the school needs signed, the birthday party you never RSVP'd to, the summer camp that's about to fill up. For any parent, it's a hundred tiny deadlines a month, and the agent catches each one and drafts the response. Things like that. Just giving some ideas to get the creative juices flowing. Tons of apps like this will be created over the next 12 months. It just makes sense.
@gregisenberg
Agent Plugins are the future of Agent Skills
@GoogleCloudTech
introducing http://omg.dev a personal computer for your coding agent. always on. work less, not more. live today.
@BennyKokMusic
We built a software factory for AI SDK. Each step is an agent, and humans merge changes. Four weeks in: ▪️ The factory authors up to 35% of merged PRs ▪️ It closed 70% of issues in July ▪️ Open bugs are down 25% https://vercel.com/blog/building-a-software-factory-for-ai-sdk
@vercel
/show-me: compact visual representations for coding agents
@dexhorthy
I made an Open Source version of the Grok Bot complete with all the features It does not need any subscriptions at all and uses the existing subscriptions you already have It can spin up virtual machines from @asciidotdev It uses @trycua for computer use It also has plugin support supporting all the integrations by @composio Routines coming soon :D Link to the repo in the comments
@milindlabs
Introducing Grok Bot, now in early beta. Bots are AI teammates that do real work for you. They sign in to your tools, use them just like you do, and come back with finished work.
@bot
Hiring Agents Is the Easy Part
@AlanaDLevin
Linux adoption among programmers is about to go parabolic. Any missing app can be recreated easily. The open-source advantage with agents is unstoppable. Nobody is going to be waiting for Apple to sign their shit. Just keep a Mac Mini as a remote builder for iOS work. Done.
@dhh
Fair warning: Omarchy is leaning fully into the future and the age of agents. This is how we truly democratize Linux, bring the malleable OS to all, and capture as much of this newfound intelligence as possible. If that's not your jam, there's a lot of anti-AI distro options ✌️
@dhh
I’ve been asked for an updated daily schedule given Superlogical, second kid, and AI usage. Here you go: - 530 awake, 20 minute interval cardio - 600 review nightly agents, start new - 700 wake and feed newborn - 730 wake, dress toddler - 745 breakfast with family - 800 off to work - 1200 lunch (often with wife) - 1530 to 1945 family time - 2000 hang with wife, start new nightly agents Early 2026 I mentioned a goal of “always have an agent running” but admitted I was realistically only getting agents for a couple hours off work. As of today, I’ve gotten a LOT better, I regularly have at least 2 agents running constantly with many running through most of the night. I’ll cover this in more detail another time!
@mitchellh
We're recording round two of this on Monday! While the calendar says it's only been a year, it feels more like twenty since we've entered the age of agents.
@dhh
Spotify has run this exact play before. In March 2020 they open sourced Backstage, the internal portal their engineers built to manage 14,000 software components. It became the industry standard for developer portals. Over 3,400 companies adopted it, including Netflix, American Airlines, and Expedia. Then Spotify built a paid enterprise product, Portal, on top of the free standard they created. Xirp is the sequel, aimed at a bigger prize. Look at what it actually does. Every agent session runs in its own git worktree, so 50+ agents can work on the same codebase in parallel without touching each other. And context lives outside the agent itself, so you can switch from Claude Code to Codex to Gemini mid-project and the full working state carries over. That second part is the whole strategy. Anthropic, OpenAI, and Google are spending billions to make their agent the one your team depends on. Xirp makes them interchangeable. Your dependency moves up a layer, to the environment holding your sessions, your service catalog, and your architectural decisions. Which happens to plug into Portal, the product Spotify sells. The model vendors are fighting to own the agent. Spotify is betting the agent becomes the commodity and the context becomes the product. Last time they made that bet, 3,400 companies took the free version and Spotify sold them the enterprise one. The music streamer from Stockholm is quietly building one of the better dev tools businesses in the industry, and almost nobody noticed the first time.
@aakashgupta
the coupling of a filesystem to agents was so dumb why did we do this if you want to support skills even for a dumb in memory agent loop you need a virtual fs why why why
@thdxr
If you’re not reading the code, whether explicitly or through agentic inquiry, one or more of these is true: ○ You’re a beginner ○ Software is throwaway ○ You’re prototyping ○ You have no users / revenue ○ You’re taking on debt & risk ○ Your problems are basic And btw. All of this is fine. But the reality is that models are still not at the “full autonomy” stage yet. They make rookie mistakes, they go down bad architectural paths. I just had the best model in the world add a nonsensical 700ms delay to “settle” something and it told me “you’re right, I was cargo-culting” 🤨 I am on the camp that this need will diminish more and more. Most code is indeed going to be assembly-like. But we also have the global internet and software infrastructure riding on these models and narrative, and we have to respect that.
@rauchg
Today, we kill the AI agent and introduce the AI employee: Lindy Teammate. It’s just like working with a real employee. Everyone on your team can simply hit it up on Slack and get 10x more done. Lindy also keeps learning, and becomes the self-updating brain of the whole business. Live now:
@Altimor
Are agents really killing UI?
@posthog
Great to see Dell collecting their well-deserved accolades for the XPS 13. And doubly so that YouTube reviewers are now focusing so much on Linux compatibility! The XPS 13 is a great choice for an Omarchy terminal/agent deck. https://www.youtube.com/watch?v=A2B7oI0FYqo
@dhh
this is great! but don't leave claudes alone :( in herdr any agent can talk to any agent. no protocol, no MCP, no setup. it's just terminals.
@herdrdev
Omarchy and AI makes it so easy to customize this. I just prompted my agent to make a personal plugin for the speedtest that "looks like a mustang speedometer", and voila, I had dhh.speedtest plugin. AI truly enables the malleable operating system!
@dhh