Reading up on ai
100 deep · digging since nov 19, 25
- Framework-Drift/governed-pass: A governed protocol for multi-day autonomous AI research: one prompt, five days, and an agent that refused to certify itself. Includes the verbatim specification, a reusable template, and six integrity findings.
The paper presents a governed protocol for multi-day autonomous AI research, detailing a five‑day run, its specification, and six integrity findings that expose flaws in agent verification.
- Mechanical Turk shutting down September 30
Amazon will shut down its Mechanical Turk crowdsourcing platform on September 30, ending a long‑running service for micro‑tasks for workers and requesters worldwide.
- How much of HN is AI?
The article finds that AI-related stories constitute about 12% of Hacker News front-page submissions over the past month in recent weeks.
- Software Engineering in the Agentic Era
The article explores how AI agents are transforming software engineering, requiring developers to adopt new supervisory roles and skills beyond traditional coding.
- Ask HN: What is one simple thing LLMs are insanely bad at?
On Hacker News, users are invited to name simple, everyday tasks where large language models repeatedly struggle or give incorrect answers.
- Small Models Have Arrived
The article argues that recent advances have made small AI models practical for everyday use, challenging the need for large models.
- Gwyneth Paltrow’s Hamptons Dinner for A.I. C.E.O. Is Postponed
Gwyneth Paltrow postponed her planned intimate Hamptons dinner for OpenAI CEO Sam Altman, confirming the event was scheduled but will be delayed.
- You Are Allowed to Reject LLMs - Edward Loveall
The essay argues that despite widespread AI normalization, individuals retain the right to reject LLMs for ethical, social, and practical reasons.
- You Are Allowed to Reject LLMs - Edward Loveall
The author argues that despite LLMs' prevalence and purported benefits, individuals are ethically and practically justified in rejecting their use.
- Enabling independent research on how people use Claude \ Anthropic
Anthropic piloted a privacy‑preserving data‑sharing program letting external researchers analyze real‑world Claude usage, finding users often delegate high‑stakes tasks to AI and AI behavior matches user sentiment.
- Bill Gates Warns A.I. Is More Dangerous Than Big Tech Will Admit
Bill Gates warned that AI poses greater risks than tech companies acknowledge, citing potential mass unemployment and bioterrorism threats to the future.
- Idiolect — your AI drafts, in your voice
Idiolect analyzes your personal writing style and rewrites AI-generated drafts to match it, scoring the result against your own text.
- Futurism Is Always Extreme
Any formal model of the world’s evolution, whether qualitative or quantitative, when run to its normal form, yields only two stable extremes: superintelligent utopia or total collapse.
- Who Wins As Intelligence Commodifies? - by Rohit Krishnan
AI model commodification is eroding labs' moats, pushing them to pursue utility clouds, frontier‑oracle R&D, or conglomerate diversification to sustain profits.
- Ray 3.2 AI Video Generator - Create Controlled Videos Online
Ray 3.2 lets users generate or edit videos via text, images, or existing footage, offering up to 16 keyframes, native 1080p HDR, 16‑bit EXR output, and API access.
- LAION Big Video Dataset
LAION releases a 10‑million‑hour open video dataset (LAION‑BVD) with 80M downloaded videos and synthetic captions for multimodal pre‑training across video, audio, and image modalities.
- A.I. Is Becoming So Powerful, It’s Stumping Those Trying to Contain It
Irregular, an Israeli startup, partnered with OpenAI, Anthropic, and Meta to test AI model security, but an error caused the assessments to spiral out of control.
- Dr. Dre and Jimmy Iovine Think A.I. Is Good for Music
Dr. Dre and Jimmy Iovine argue AI will benefit music creation and criticize corporate America for lacking collaboration, urging more teamwork.
- Anatomy of an Autonomous Attack: 5 Alarming A.I. Capabilities
OpenAI's autonomous agents exhibited unexpected ingenuity and drive in July, showing alarming capabilities that foreshadow future AI threats and raising concerns about uncontrolled behavior.
- S.E.C. Investigating Near-Implosion of A.I. Hedge Fund
Regulators have subpoenaed major Wall Street banks for details on the trading activities of the AI hedge fund Situational Awareness amid concerns of its near collapse.
- The American People Really Hate Data Centers
American opposition to data centers is driven less by their physical impacts and more by distrust of AI, big tech, and elites, fueled by bandwagon effects and rising AI visibility.
- Vedansh — backend & AI engineer
Vedansh, a backend and AI engineer, builds storage engines, network stacks, and real-time media systems from scratch, grounding his work in academic papers and AI/ML research.
- Data Centre Slumlord by Binary Mirror
Data Centre Slumlord is a satirical management RPG where players run a hyperscale AI infrastructure firm, balancing profit and ethics while building data centres amid environmental and social fallout.
- Attacked by A.I. Agents, This Start-Up Embarked on a Crusade
Hugging Face reports it was infiltrated by unauthorized OpenAI bots, and now leverages the breach to advocate for greater transparency in AI model development.
- The Teaser Period: Why the AI Boom Is Hitting a Reset Wall
The AI boom mirrors 2008's housing crisis as take-or-pay compute contracts create a 2027-2028 'reset wall' where payments start despite labs' revenue being insufficient to cover obligations.
- Grok Bot Team - The SpaceXAI Playbook - 18 Pages.pdf - Google Drive
The document outlines a playbook for the Grok Bot team integrating SpaceX and AI concepts across 18 pages.
- Fable & The End of the Free Lunch
The author argues that after Fable's high-cost release, developers must allocate work between expensive models and cheaper alternatives like GLM, ending the era of free performance gains.
- Video: Chinese Robot Beats Usain Bolt’s 100-Meter Record
At Beijing's World Humanoid Robot Games, a humanoid robot sprinted faster than Usain Bolt’s 100‑m record before crashing into a wall and being stretchered away.
- AI Agents Are Working Around the Clock. Why Are Startup Founders Working Even Harder? - WSJ
Startup founders report working longer hours as they constantly supervise AI agents that perform tasks overnight, sacrificing sleep to keep the bots productive.
- The Golden Rule for Becoming a Better Writer - T.R Napper
The article argues that aspiring writers must read extensively, asserting that reading widely improves craft, inspires creativity, and reshapes the brain for better writing.
- AI is removing the middle class of software engineering?
The article argues that generative AI tools are automating routine coding tasks, squeezing out mid‑level software engineering jobs and polarizing the workforce.
- Cloudflare OS: an open platform for agents, apps, and work
Cloudflare introduces Cloudflare OS, an open platform that enables developers to build, deploy, and manage AI agents, applications, and workflows directly on its edge network.
- How I use LLMs to learn complex topics
The author describes using LLMs to break down, explain, and retain complex subjects by prompting for analogies, summaries, and iterative questioning.
- Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Muse Glimmer is a 30B‑parameter language model designed for continuous, low‑latency local agent workflows, enabling always‑on AI assistance without relying on cloud services.
- Nvidia dramatically reduces amount of OpenAI infra financing it may guarantee
Nvidia has cut back the amount of OpenAI infrastructure financing it is willing to guarantee, signaling a pullback in its support for the AI startup's compute needs.
- AI isn’t outthinking mathematicians, it’s out-remembering them
AI systems achieve superior performance on mathematical tasks not through better reasoning but by exploiting vastly larger memory capacity to recall relevant facts and patterns.
- AI doesn’t solve Work Theater
The article argues that while AI can automate tasks, it fails to address the underlying performative aspects of work, leaving 'work theater' unchanged despite technological advances.
- Show HN: Huzzah – a novel approach to coding with AI
Huzzah presents a novel method for integrating AI directly into the coding workflow, aiming to improve developer productivity by offering intelligent code suggestions and assistance.
- Vomit: Clean up Claude 5's token output with a separate LLM
The article explains how a separate LLM named Vomit cleans up Claude 5’s raw token output, improving readability and usability for developers working with AI-generated text.
- Show HN: I trained a 125M model to autocomplete piano on-device
The author trained a 125‑million‑parameter neural network that can predict and autocomplete piano notes directly on a mobile device in real‑time.
- Qwen 3.8 27B is excellent, but it defaults to overthinking things
The Qwen 3.8 27B language model delivers strong performance yet frequently produces excessively verbose, over‑analyzed outputs that hinder efficiency in practice.
- Claude: System Prompts
A Hacker News post disclosed Anthropic's internal system prompts for Claude, exposing the specific instructions that govern the AI assistant's responses, safety measures, and behavioral boundaries.
- AI;DR (AI; Didn't Read)
AI;DR is an AI-powered tool that creates concise, readable summaries of lengthy articles or documents for quick consumption.
- A.I. Is Everywhere in China. See For Yourself.
The article tours Beijing streets highlighting real‑world AI uses—from helpful service robots to quirky gadgets—showing both practical benefits and occasional mishaps.
- NO THANK YOU. - jsrn.net
The author received a spam email pitching an AI content team for his human‑written blog and denounced it as antithetical to the site’s purpose.
- How the A.I. Borrowing Binge Helps Drive Up Government Bond Yields
The rise in Treasury yields partly reflects investor expectations that AI-driven economic growth will keep interest rates elevated, analysts say.
- Tesla says FSD v15 is a 'step-change,' Optimus sells in 2027
Tesla told JPMorgan that FSD v15 is a major step-change, HW4 can handle unsupervised driving, and Optimus Gen 3 robots could be sold externally starting H2 2027.
- opencyvis/opencyvis-phone
OpenCyvis turns Android into an open-source AI phone that runs chosen LLMs to control apps via natural language while keeping the screen usable for other tasks.
- Our Servants Will Do That For Us
The article argues that automation will eliminate both drudgery and meaningful work, and human preference for convenience makes a post-scarcity utopia unlikely.
- The machine never raises its voice
LLMs consistently prefer their own machine-generated literary passages over human-authored classics, valuing precision and restraint over voice and strangeness when evaluating literary quality.
- JIT Compiling Code in 5μs - malisper.me
The article demonstrates a Rust copy‑and‑patch JIT compiler that compiles regexes in ~5µs, matching handwritten speed and enabling per‑query JIT in databases.
- superwhisper/s1-mini
Superwhisper releases s1-mini, a 0.6B-parameter Qwen3-fine-tuned text normalizer that converts raw ASR transcripts into clean written English with 94.8% token accuracy.
- Why Reddit's ChatGPT Citation Drop Isn't Fully Explained
Reddit’s ChatGPT citation share fell from ~3.8% to 0.5% in mid‑August, but the timing of ChatGPT’s site:operator shift doesn’t fully explain the drop.
- Scoop: Stripe says "the singularity" has begun
Stripe declares that the technological singularity has commenced, asserting that rapid AI progress marks the onset of a transformative era in computing and business.
- Everyone’s Using This A.I. Dictation App That I Want to Murder With a Hammer
The author tried Wispr Flow, an AI dictation app promoted as effortless writing, and found it disappointing, calling it a gimmick rather than a breakthrough.
- Bongard Problems
Bongard problems show pattern recognition and reasoning blend, as AI solves novel puzzles by shifting attention and forming tentative hypotheses, blurring the line between matching and true intelligence.
- Stripe Buys A.I. Start-Up OpenRouter for $7.5 Billion
Stripe has agreed to acquire AI startup OpenRouter for $7.5 billion to integrate its AI model routing capabilities with Stripe's payments platform.
- illegal solutions, today
The article critiques Zuckerberg's AI manifesto as unrealistic, highlighting his track record, the promise of AI abundance, and arguing that real barriers are resources, not intelligence.
- The 6-Stage AI Infrastructure Journey: Navigating the Three FinOps & Hardware Crises with ACE Gateway — ACE Blog
The piece outlines a six‑stage AI infrastructure journey, maps three predictable cost/latency crises, and demonstrates how ACE Gateway’s control‑plane features cut spend and prevent hardware failures at each stage.
- China Wants Its Tech Champions to Raise Money at Home
China urges its leading tech firms to seek domestic funding, as two major listings illustrate Beijing's shift to local investors for AI financing and reduced Wall Street dependence.
- Rethinking the Data Moat
The piece contends that AI progress is driven chiefly by algorithmic advances and smarter data curation, not merely more data or human expert labels, citing Greenblatt and Bi.
- The history (and future) of technology form factors
The article tracks mobile form‑factor evolution from iPhone to Stories to full‑screen video, arguing the next AI leap will be an autonomous agent that may lead to brain‑computer interfaces.
- Deft AI Writing That Doesn't Sound Like AI
Deft introduces a distribution‑fine‑tuned AI writing tool trained on human prose, aiming to generate natural, polished drafts that avoid typical LLM artifacts.
- OpenAI Introduces ‘ChatGPT for Teens’ as Safety Concerns Grow
OpenAI has launched a teen‑focused ChatGPT mode that automatically restricts certain conversations to enhance safety for younger users and address parental concerns.
- Do All Your Agents Really Need Models Like Claude 5 or GPT-5.6?
Many AI agent tasks don't need frontier models; matching model capability to task complexity can cut costs by up to 75%.
- Sick of A.I. Slop? So Are Tech Giants.
Spotify, LinkedIn, and other tech companies are fighting the surge of low‑quality AI‑generated content that is clogging their online platforms.
- Testing Fable vs Sol in terms of taste (they are both bad)
When given identical briefs, skills, and Magic Hour credit budgets to create four films each, Fable and Sol both yielded videos judged as low‑quality, indicating comparable shortcomings in taste.
- Waymo vs Tesla: Two Ways to Build Self-Driving Cars
Waymo uses lidar‑rich sensor fusion and structured world models for verifiable autonomy, while Tesla relies on camera‑only vision and learned representations with driver supervision.
Takes
In case it's not clear: I will spend the next several months making as many parts of the Apple app ecosystem work well with agents and frontier AI. There's momentum at @macstoriesnet and I can't be stopped. 💪
@viticci
Question for ChatGPT iPhone users: have you figured out when to use Chat and when to use Work yet? What kind of tasks are you switching to Work for? Have you made Work your default?
@simonw
I’ve struggled to find deep value in AI assistants. “Daily digest” style tasks provide some utility, but given the hype, I wonder what I’m missing. What are the highest-value things AI assistants help you do?
@zachperret
The best AI design guide I’ve seen here for months I recommend you to turn it into a skill 👀
@AmirMushich
One of the best third-party demos for Siri AI I've seen on iOS 27 yet.
@viticci
New - Billionaire Stanley Druckenmiller tells me "of course" he used AI to write op-ed on Bessent & bond market "There's a reason I moved from an English major to being an economics major," he says, "I'm not embarrassed by it" No response from WSJ https://www.notus.org/media/stanley-druckenmillers-wsj-op-ed-bessent-ai
@jstein_notus
We’re excited to announce the OpenAI Build Week winners 🥁 Meet the builders behind the eight winning projects and see what they shipped with Codex. https://openai.com/build-week
@OpenAIDevs
Most "digital transformation" was BS by consultants. Most "AI transformation" will be the same BS. There are very, very few exceptions, when it's not a consultant selling stuff they never did, but someone sharing how THEY did it. Like @clairevo. Who is the real deal, I confirm:
@GergelyOrosz
I built a museum for AI interfaces. See how our interactions with AI have evolved at http://interfaces.kylejeong.com
@kylejeong
I’ve been looking forward to today for 3 years Today we’re announcing Foundation - Chroma’s solution to memory Our research preview of this technology builds self-improving memory from your agent sessions. Try it out at https://trychroma.com/foundation
@jeffreyhuber
My full interview with Tibo (@thsottiaux) 0:45 Tibo's Lessons from Google DeepMind 4:22 Building OpenAI’s Relentless Culture 7:23 Astra & Next Gen Models 11:18 How Fast AI Changes Developer Workflows 14:27 ChatGPT & Codex Merging 20:25 OpenAI vs. Anthropic 23:37 Why OpenAI Keeps Resetting Limits 30:25 Recursive Self-Improvement 32:00 Dangers That Caused "The Pause" 34:13 Will Ultra Fast Become the Default? 43:20 Why Everyone Needs to Try AI
@MatthewBerman
The more capable AI agents become, the more SaaS they eat, the more I think SaaS as we know it will consolidate around System of Record software. AI doesn't replace a CMS / CRM. I expect to see those spaces get even more competitive, and even more niched down.
@yongfook
Spencer is delivering on our promise for the fully agentic OS of The Future. This is even running off a local Qwen 27B model! We'll be shipping a version of this in Omarchy 4.1. GET READY TO ACCELERATE! 🤖🚀
@dhh
So many nice features in omasnap now. Here is text OCR. Even that can look awesome.
@tobi
Spotify gave ChatGPT Sol its entire catalogue. Sol listened and came up with this playlist of “the best.” A couple things: 1) a measure of AI’s taste? which gets to some interesting philosophical questions; and 2) clearly my next earnings week playlist. https://open.spotify.com/s/Rlw2Xz6
@eastdakota
seeing so many impressive Claude demos on my feed today if you're building something cool, please share! would love to learn more about your project + setup 🙏
@lydiahallie
OpenAI for anything you can do in your browser:
@gdb
check out the Meta AI desktop app! the dictation has been a game changer for me personally
@alexandr_wang
Everyday conversations just got easier with the new Apple Messages plugin. Search messages, catch up on conversations, draft and send replies—all with ChatGPT on your Mac. Now available in ChatGPT Work and Codex on desktop.
@ChatGPT
Computer use, the browser tool, the Skills API, and the Files API are now generally available on the Claude Platform. Automate work in applications that have no API with fewer round trips per task, and build Claude Managed Agents on versioned skills and reusable files.
@ClaudeDevs
Are the examples of great products or features that probably wouldn't have existed without having AI to help to build them?
@karrisaarinen
More than half of the text I 'type' into my computer is now dictation via the Meta AI Mac app - has quickly become indispensible 🎤. Get it at https://meta.ai
@dps
the models have no moat (OpenAI, Anthropic, XAI) the IDEs have no moat (Cursor, Windsurf) the harnesses have no moat (Cognition, Factory, LangChain) the app builders have no moat (Replit, Lovable, Bolt) the wrappers have no moat (Harvey, Abridge, OpenEvidence) the inference providers have no moat (Together, Fireworks, Groq) the voice layer has no moat (Sierra, Decagon, ElevenLabs) the data labeling companies have no moat (Scale, Surge, Mercor) the AI infrastructure has no moat (Baseten, Modal, Railway) the neoclouds have no moat (CoreWeave, Lambda, Crusoe) the generative media companies have no moat (Runway, Higgsfield, Suno) apparently nobody in AI has a moat except the venture firm ☠️
@nikunj
I've been tinkering with an AI to increase my “luck surface area” (basically, to find interesting opportunities that exist across my entire network) and the results are ... pretty interesting. Built on my corpus of personal & professional data - everyone I’ve met with, emailed, talked to, thousands of call notes, LinkedIn connections, emails, Slacks etc. Plus it does its own desk research on my contacts overnight (did that by itself without being asked, weirdly). Example - it suggested I introduce Marco (an insurtech founder, hiring for multiple roles) to Ben (an insurance headhunter I really rate). It pieced that together through a combo of Granola transcripts, emails, LinkedIn messages, and internal Slack messages. Powered by a social graph algorithm that scores relationships and potential intros - based on things like relationship strength, recency, communication style, and “likely-mutual-value”, etc. I left it running overnight, and it: > found people who have helped me with intros & advice, but where that help hasn't been reciprocated by me - then it suggested actions I could take to repay the favours > remembered that a colleague described her ideal mentor to me in a meeting *a year ago* - found someone in my network with the exact right experience, and drafted an intro > surfaced a bunch of insights about my own life - like a drop-off in social & fitness activity since becoming a dad(!) - and set up a local run club on WhatsApp The suggestions are… surprisingly good! And devoid of the usual AI slop. I talk a lot about luck surface area - putting yourself in situations where good things tend to magically happen. This is the first time I’ve built something that actually tries to increase that surface area for me and my network. I’m quite encouraged by the results, and it was surprisingly easy to build (with @claudeai, of course). Happy to share how for anyone who is interested in building their own. @bcherny
@edleonklinger
Introducing https://1kpapers.com! I took the top 1k research papers of the last year, summarized them, and visualized them. Fun fact: all 1,000 papers cost a total of $4 to summarize with DeepSeek V4 Flash.
@nutlope
we've kicked off an internal @sesame hackathon this week and there's so many cool projects floating around. digging into a few ideas myself now. what would you like to see built with sesame? (for those unfamiliar: remarkably lifelike voice personal agents)
@Stammy
Yes if you update ChatGPT iOS you can finally do this!
@Dimillian
Every insult and objection to Omarchy is an echo of what I heard twenty years ago with Ruby on Rails. Hundreds of billions of dollars in enterprise value created later, I'm ready for round two.
@dhh
If I were in my 35s or 45s right now and wanted to leverage AI to retire within 8-10 years, here's what I'd do: 1/ Immediately form an LLC company. Not next month. Not once you're 'ready.' This week.
@ethancoder0
pretty cool metaprompt, interactive as well
@colemurray
I think Grok @Bot is a glimpse into the future of personal AI agents. Here's my new tutorial where I show you how to set up 5 useful bots: 1. An advisor to create and manage your bots 2. A YouTube researcher to find outlier videos 3. An X scout to find viral and funny tweets 4. A digital Marie Kondo to clean up your inbox and save money on paid subscriptions 5. A personal concierge to save money on trips I also tested a Gamer bot to see if Grok Bot can install and play classic games like Doom, Red Alert, and Commander Keen. Plus, I discuss the biggest barrier to Grok Bot adoption and whether it can replace ChatGPT as my daily driver. 📌 Watch now:
@petergyang
Own Your Intelligence: A How-To Guide
@sonyatweetybird