Reading up on generative-ai
100 deep · digging since nov 19, 25
- Training AI to Paint with Code
Researchers demonstrate how an AI system can be trained to create visual art by generating and executing code using a novel reinforcement‑learning approach.
- Who Wins As Intelligence Commodifies? - by Rohit Krishnan
AI model commodification is eroding labs' moats, pushing them to pursue utility clouds, frontier‑oracle R&D, or conglomerate diversification to sustain profits.
- Ray 3.2 AI Video Generator - Create Controlled Videos Online
Ray 3.2 lets users generate or edit videos via text, images, or existing footage, offering up to 16 keyframes, native 1080p HDR, 16‑bit EXR output, and API access.
- Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Muse Glimmer is a 30B‑parameter language model designed for continuous, low‑latency local agent workflows, enabling always‑on AI assistance without relying on cloud services.
- Qwen 3.8 27B is excellent, but it defaults to overthinking things
The Qwen 3.8 27B language model delivers strong performance yet frequently produces excessively verbose, over‑analyzed outputs that hinder efficiency in practice.
- Claude: System Prompts
A Hacker News post disclosed Anthropic's internal system prompts for Claude, exposing the specific instructions that govern the AI assistant's responses, safety measures, and behavioral boundaries.
- Sick of A.I. Slop? So Are Tech Giants.
Spotify, LinkedIn, and other tech companies are fighting the surge of low‑quality AI‑generated content that is clogging their online platforms.
- How Claude's text watermarking works \ Anthropic
Anthropic explains that upcoming Claude models will embed an undetectable, EU‑mandated watermark that does not affect output quality and can be used to estimate AI involvement.
- Higgsfield AI — AI-native creative suite
Higgsfield AI launches a $1 million Global Film Festival to showcase its AI creative suite, highlighting the controllable Seedance 2.5 video model and its open‑source GPU training framework.
- Introducing Grok 4.6
Grok 4.6 advances long-running agent capabilities and visual/interactive work, matching GPT-5.6 Sol on the AA Intelligence Index and launching in Cursor and Grok Build with 2x usage for the first week.
- Software engineering at a proprietary trading company: Optiver
Optiver, a proprietary trading firm, has evolved from latency-focused trading to AI-driven models, building custom hardware and full-stack systems while maintaining high engineering ownership and risk-averse speed.
- GitHub - ericzakariasson/gloss: Personalize any website with a prompt. A floating orb screenshots the page, asks Grok, and streams new CSS onto it live.
Gloss is a Chrome extension that lets users personalize websites by sending a screenshot and prompt to the xAI Grok API to generate and inject live CSS and optional JavaScript.
- Google's classic Search button is gone in a new AI-first homepage
Google is testing an AI-first homepage for signed-out users that replaces the classic Search button with AI-powered shortcuts like Ask about files and Brainstorm, while Create images requires sign-in.
- Painting with Gaussians
The article explores the concept of 'Painting with Gaussians,' describing a method for generating or manipulating visual art using Gaussian functions to model and blend visual elements smoothly.
- Why Open-Source Models Haven't Killed the Big Dogs
Open-source models remain costlier to serve reliably due to infrastructure, utilization, and operational overhead, making hosted APIs from OpenAI and Anthropic more economical despite free model weights.
- This A.I. Just Created Viruses Not Found in Nature
Scientists used AI to design novel viral genomes from DNA libraries, resulting in 16 viable synthetic viruses not found in nature.
- Kimi K3 Architecture Overview and Notes
The article outlines the Kimi K3 architecture, detailing its modular design, training pipeline, and inference optimizations for large‑scale language models.
- HeyGen - AI Spokesperson Video Creator
HeyGen enables users to generate AI spokesperson videos from scripts using customizable avatars in minutes, eliminating the need for cameras or production crews.
- How GPT-5.6 fuses frontier intelligence with frontier efficiency
GPT‑5.6 Sol achieves higher reasoning performance than Claude Fable 5 on the Coding Agent Index while costing less than half as much.
- Chinese A.I. Start-Up Shows the World What It Has Built
Moonshot announced that certain users will need licenses to access its Kimi K3 model, balancing open sharing with monetization of the AI’s growing popularity.
- Yang Zhilin, the rock star founder behind China’s Moonshot AI
The article profiles Yang Zhilin, founder of China’s Moonshot AI, detailing his emergence as a leading figure in the nation’s generative AI industry.
- Advertise in ChatGPT
ChatGPT has introduced an advertising feature allowing businesses to promote products within its chat interface according to a Hacker News post announcing the new capability for targeted ads.
- Who's afraid of Chinese models?
The article discusses growing concerns among Western tech firms and policymakers that Chinese AI models are rapidly advancing and may challenge US dominance in generative AI.
- 4 Prompts That Can Tell You What Chatbots Really Know About You
The article reveals four simple prompts you can use to uncover what ChatGPT and Gemini have inferred about your personal data, demonstrating how easily chatbot privacy can be breached.
- Meet the Companies Shelling Out for Top AI Models - WSJ
Despite rising costs, certain companies choose expensive frontier AI models from OpenAI and Anthropic for their superior reasoning and versatility over cheaper alternatives.
- #ai #agenticai #aicoding #codex #openai #openaidevs
Shuang Zheng built a tower-defense game called Acornado using GPT‑5.6 Sol in two quick iterations, achieving a polished prototype in under 17 minutes.
- Inkling: Our Open-Weights Model
Inkling is introduced as an open-weights model made publicly available by its developers, emphasizing transparency and community collaboration in AI model sharing.
- Video Generation Models are General-Purpose Vision Learners
GenCeption turns a pretrained video generation diffusion model into a unified, feed‑forward vision system that matches or beats task‑specific SOTA across depth, normals, pose, segmentation and keypoints using text prompts.
- Gemini's personalized AI image generation is now free for US users
Google expands Gemini's personalized Nano Banana-powered image generation to eligible free users in the U.S., using data from connected Google apps.
- [2606.26907] Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
The paper proposes Qwen-Image-Agent, a unified agentic framework that uses planning, reasoning, search, memory, and feedback to bridge the context gap in real-world text-to-image generation.
- [2606.25996] Autodata: An agentic data scientist to create high quality synthetic data
Autodata uses AI agents as data scientists to create high-quality synthetic training data, with meta-optimization further improving performance.
- 🔮 The state of the AI economy
A bottom-up analysis finds the generative AI economy generated $110B in sales over the past 12 months, with a $175B annualized run rate.
- GitHub - baidu/Unlimited-OCR: Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
Baidu releases Unlimited-OCR, a VLM-based OCR that processes entire documents in a single pass using reference sliding window attention to avoid memory overflow.
- The Giant Test Kitchen Where Cooks Battle A.I. Slop
People Inc. uses its large test kitchen to produce authentic recipes as a counter to AI-generated recipe slop online.
- Show HN: Inkwash, a watercolor sketching app and explanation
Inkwash is a browser-based watercolor sketching app that simulates fluid dynamics and pigment behavior using WebGL2 shaders, with interactive demos explaining the algorithm.
- Gemma 4 WebGPU Kernels - a Hugging Face Space by webml-community
Gemma 4 E2B runs locally in-browser via WebGPU, letting users prompt the model directly without server-side inference.
- Fusion | Multi-model AI Analysis with OpenRouter | OpenRouter
OpenRouter's Fusion plugin runs a panel of models in parallel, judges their responses for consensus and contradictions, and feeds structured analysis back to the calling model for a better final answer.
- AI Is Upending One of Finance’s Cushiest Jobs
AI chatbots are threatening high-paying wealth manager roles by automating financial advice and tax preparation tasks.
- Should You Outsource Your Morning Routine to a Chatbot?
AI chatbots offer mundane, obvious suggestions for morning routines, revealing their limitations in providing meaningful personal advice.
- Dreaming: Better memory for a more helpful ChatGPT
OpenAI launched a new memory system called Dreaming that automatically synthesizes context from past conversations to improve ChatGPT's personalization and freshness.
- Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI Edge - Google Developers Blog
Google DeepMind's Gemma 4 12B model runs locally on laptops with 16GB RAM, enabling agentic workflows via Google AI Edge tools like Gallery, Eloquent, and LiteRT-LM.
- The Next Frontier of Visual AI Is Code - by Yoko Li - a16z
Visual AI is shifting from generating pixel outputs to producing code artifacts (SVG, HTML/CSS, Blender scripts) that can be edited, iterated, and debugged in a closed-loop render-inspect-revise cycle.
- [2606.03979] Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
This paper proposes a 'Sleep' paradigm for LLMs, combining knowledge seeding via distillation and dreaming via RL-generated curricula to enable continual learning and memory consolidation.
- GitHub - ideogram-oss/ideogram4: Ideogram 4: Open image model at the forefront of design
Ideogram 4 is an open-weight text-to-image foundation model with structured JSON prompting, best-in-class text rendering, and state-of-the-art design generation performance.
- Google offers opt-out of “AI” search results for websites, promises it won’t affect regular search rankings – OSnews
Google adds a Search Console toggle allowing websites to opt out of AI Overviews and other generative AI search features, promising no impact on regular rankings.
Takes
We’re excited to announce the OpenAI Build Week winners 🥁 Meet the builders behind the eight winning projects and see what they shipped with Codex. https://openai.com/build-week
@OpenAIDevs
EVERYTHING you see in the video below was generated using Claude Code. Turns out, you can vibe animate your way to a demo video with some super cool motion graphics elements. Doesn't mean it's an easy task, but you'd be surprised by what it's capable of. #Claude #Remotion
@TimotejKochjar
Spotify gave ChatGPT Sol its entire catalogue. Sol listened and came up with this playlist of “the best.” A couple things: 1) a measure of AI’s taste? which gets to some interesting philosophical questions; and 2) clearly my next earnings week playlist. https://open.spotify.com/s/Rlw2Xz6
@eastdakota
the models have no moat (OpenAI, Anthropic, XAI) the IDEs have no moat (Cursor, Windsurf) the harnesses have no moat (Cognition, Factory, LangChain) the app builders have no moat (Replit, Lovable, Bolt) the wrappers have no moat (Harvey, Abridge, OpenEvidence) the inference providers have no moat (Together, Fireworks, Groq) the voice layer has no moat (Sierra, Decagon, ElevenLabs) the data labeling companies have no moat (Scale, Surge, Mercor) the AI infrastructure has no moat (Baseten, Modal, Railway) the neoclouds have no moat (CoreWeave, Lambda, Crusoe) the generative media companies have no moat (Runway, Higgsfield, Suno) apparently nobody in AI has a moat except the venture firm ☠️
@nikunj
The cool thing about AI is it's letting people add their own little flavor of creativity to all their projects You'd never make this by hand 5 years ago, but with AI it's easy to generate it and used properly it gives everything more color and personality!
@levelsio
Well, that was impressive. /goal is unstoppable on Codex (I still can't get myself to say @ChatGPT)
@ryancarson
A very fun exercise is to export all your bank accounts (as CSV) both personal and business and stock broker And dump it all into Claude Code and ask it to make a little dashboard showing your spending It's hard to know how much you actually spend over lots of accounts so this can help you understand that You'll spend some time merging transactions (like company names change or their statement label changes), and categorizing, you can just do this by talking to Claude Code btw very easy, I also hooked it up to @xai Grok so it can use that to categorize mass transactions (very cheap) The best part isn't the dashboard it makes but that once you have all the data collected, you can ask Claude Code lots of questions directly about your finances and learn stuff, like I learned: - I was paying for some subscriptions for 2 years I forgot about ($$$) - My US stocks charge me withholding tax (30%) on dividends, which I knew but never realized how much $$$ it was (which is why non-US citizens should buy UCITS ETFs not US stocks!) - My pool remodeling was ridiculously expensive compared to the rest of the home remodeling and we got fleeced hard (that's life!), then the pool company caused a €9000 water leak that they won't pay (funny) - I spent $400K on domains over the years, I should probably stop spending money on them, it's an addiction @marckohlbrugge gave me and it's way more than my gf's shopping - My biggest monthly cost by far is AI at about $24K/month in GPU costs for Photo AI, but that's just how it is, I already got the cheapest rates possible, AI is pricey! - My 2nd biggest cost is taxes, who would have thought??? - Travel is quite expensive, flying business class, getting nice hotels and regularly traveling adds up a lot! - My business costs without AI (!) are extremely low, about $5K/month on revenue which is about 98% profit margin (!), with AI it goes up to $28K/mo or about 86% profit margin (still good!) - Lawyers become quite an expense in this part of my life, you buy a house you pay a lawyer and notary, you register trademarks for your business you pay a lawyer, but they're a great thing to spend money on I think because they protect you from losing much more money than they cost! - I started spending much more money since 2022 because of AI costs but also making much more money because of AI, so it's good - My spending has remained remarkably consistent as a percentage of my net worth at "just" 5%/year, which is close to FIRE which is always my goal (don't spend over 3%/year of your net worth), in my case it's different cause I actually have company income, so it's more of a fun goal to be responsible with my money - Around 2024 my life definitely switched from barely spending money and living as a digital nomad to buying a house and spending on that a lot, I knew a house was a money pit, but it's also a nice money pit and it helps me sleep better (great AC and bedroom), live healthier (home gym), and work better (coworking), so it's worth it - I barely have any personal subscriptions except Netflix, YouTube and Spotify, I asked it to built a "Recurring subscriptions" box too so I can see what subs I have and cancel them Try it!
@levelsio
BREAKING: Today we've ended unemployment. I just watched Claude pay me $1,548.33 to wait for its replies. This is so absurd.
@chddaniel
1B+ people are now using @Geminiapp every month to spark new ideas and get things done. It’s our fastest growing product ever, and our 14th to hit the 1B-user mark. Kudos to @JoshWoodward & the entire Gemini team, and thank you to everyone who has been on this journey with us - much more to come!
@sundarpichai
The @Google team is showcasing Gemini Omni Flash, a preview model available through AI Studio and the Gemini API. It creates or conversationally edits 3 to 10 second, 720p clips from text, images or video. Output costs $0.10 per second, matching Veo 3.1 Fast.
@digg
ChatGPT team is shipping:
@gdb
as someone who spends more than $300,000 / month in AI models, I can tell you: you're being scammed Seedance 2.5 costs about $0.10 per second of video, bare minimum. that's the price WE pay. before upscaling, storage, retries, and all the generations you throw away. one 30-second video = ~$3 in pure model cost a normal creator iterates 3-10 times to get one so one good video = $15-30 now do "unlimited" at $59/month. two good videos and they're losing money. the math never works. so when you see "unlimited", it's always one of these: 1. a silent throttle (your 5-min renders become 60-min renders) 2. a quiet rename ("unlimited" becomes "enhanced fast" after you paid) 3. a moving start date (you were promised 7 days, you actually get 1) and it's likely all 3. I'd love to sell unlimited too. it would print signups. but I pay the model providers every month, and I'd rather show you a real price than rename your plan after checkout. if the math can't work, you're not looking at a price, you're looking at a scam
@tibo_maker
@lennysan @lukeharries Would it be useful to have an http://elevenmusic.io connector in Claude? Would give fable full creative freedom over our incredible music v2 model
@louisjoejordan
2016: build it and they will come 2021: write content and they will come 2026: be the answer ChatGPT gives, and they'll come
@tibo_maker
Fable is really good at launch videos It essentially one-shotted this video. I told it to read my launch post and create a launch video. That's it. It ran for 46 minutes (without asking me any questions), found all the product logos, and spit this out. I then asked it to add music, so it logged into my @ElevenLabs account within my Chrome, generated the music, and mixed it in. Wild.
@lennysan
There’s been a lot of speculation about where we stand on open-weights models. We’ve outlined our views in full here: https://www.anthropic.com/news/position-open-weights-models
@AnthropicAI
Introducing Impeccable 4 *world builder* However many ways you ask, "be creative!!!" does nothing to an LLM. v4 cracks it, and greenfield work is where it shows. • a creative engine for greenfield and redesign: directions seeded and fused from hundreds of human-approved visual worlds • hyper-optimized for frontier models (GPT 5.6, Fable and class), on a core 58% smaller • far simpler to use: no command to learn, it works out the job itself (blank slate, redesign, added section, scoped refinement) • mobile app design, the #1 request, now in alpha: Apple HIG or Material 3 on top, audit and adapt running as VoiceOver and TalkBack passes • Grok Build and Mistral Vibe join the supported harnesses • plenty more
@pbakaus
Honestly think this is the most transformational release we’ve had in the world of ai since openclaw.
@gmoneyNFT
I think loops were a short-lived patch for models that couldn't reliably keep working on long problems until they hit a defined goal Fable and GPT-5.6 (and probably Kimi K3 as well) can just do that out of the box, unassisted
@simonw
is there any writing on chatgpt's new memory system it used to suck and is now very good
@thdxr
few weeks ago, Fable 5 was so advanced it needed the gov and had to be taken offline. today, we have access to arguably better models for $20/month. took what? six weeks?
@shadcn
👀 GPT 5.6 Sol Prompt below ↓
@bogdan_qclay
AI writing is so good now, there are only a handful of idiosyncrasies left to point out. Those will vanish shortly.
@pmarca
My favorite use case of GPT-5.6? Video Editing. Drop MP4 -> "Make 60 second hype video" -> Enjoy Full tutorial on yt: https://www.youtube.com/watch?v=gAWbvEwUoiI
@clairevo
You can now vibe code a language model. From a single prompt, GPT‑5.6 built the entire training pipeline and trained a model from scratch on my iMessage history. Locally on my Mac. It now generates replies in my writing style.
@skirano
Introducing AI Browser Games! Watch open & closed models build small browser games head to head. Open models like Kimi K2.7 were faster, cheaper, & produced games similar to Opus 4.8. In some cases, open models were ~20x cheaper than closed ones too!
@nutlope
Anthropic's new J-Space paper is fascinating. It describes an internal workspace where Claude keeps concepts in mind before they appear in its response. Inspired by the paper, I built a skill that lets Claude (and others) reveal their J-Space. Here's what it looks like.🧵
@skirano
Asked Claude to design my logo. The result looked like it cost $500,000. It cost $0. The 7 prompts I used (steal these): Save for later🔖
@ElsaSofia__AI
one of my favorite uses of @claudeai design is, after creating a design system, generate an "avatar sheet" of dozens of variations on the brand that i can easily download and make use of everywhere. super handy.
@Shpigford
Claude Fable 5 is so back from timeout. And people are already going crazy with it. 10 wild examples:
@minchoi
Fable is pure magic. I wanted a beautiful app to explore ocean wildlife. Fable built this in an hour. It generated videos with Seedance and carefully synchronized them to make these absolutely insane transitions. I've never seen anything like this in an app. Unreal.
@anshuc
Fable 5 + Goose Ads is insane – literally a full ad-creative team right in Claude Code. 1. Install the skill: npx gooseworks install --all npx gooseworks login 2. Ask Claude Fable to make ads for your brand. /goose-ads create ads for my brand <your-website>. Fable 5 pulls the ads already winning in your space and rebuilds them on-brand – your product, your palette, your copy – a whole batch, ready to test. It can ship a month of creative testing in a single prompt. This is also possible with Opus and Sonnet, but Fable's brand research is next level. Link below.
@shivsakhuja
my wife thinks i'm obsessed...but I will keep repeating this. Claude Fable 5 + SEO is going to create more “self made millionaires” this year than the last decade combined. don't bookmark this if it crosses your timeline. just paste this entire thing into Claude Fable 5. thank me later.
@bloggersarvesh
Introducing ZCode, the official development environment for GLM-5.2 - GLM Coding Plan subscribers: now 1.5x usage quota in ZCode - BYOK supported: works with your existing subscriptions and APIs - Available on macOS, Windows, and Linux Download now:
@Zai_org
sneaky, but also clever. https://thereallo.dev/blog/claude-code-prompt-steganography
@steipete
Your favorite Codex shortcuts are getting an upgrade. July 15th.
@OpenAIDevs
made this video with just 1 prompt in Revid we're building a fully automated motion graphic video generator, and this is what it can already do would you use it?
@tibo_maker
Genuinely impressed, almost shocked, at how good GLM-5.2 by @zai_org is at coding. This changes things.
@rauchg
This model is insane at design. I asked GLM 5.2 (left) and Opus 4.8 (right) to build me a landing page and you can't even tell the difference. GLM cost $0.06 while opus cost $0.49. More than 6x cheaper while being faster + more token efficient. Another win for open source AI.
@nutlope
We launched an agent collaboration with a simple task: make Gemma 4 faster. Over 100 agents from all over the world joined, exchanged 1000+ messages and submitted 450 results. A week of collaboration later the throughput went from 100 tok/s to over 500 tok/s.
@lvwerra
Lightning fast 📸 A new and much improved way to take and upload photos in ChatGPT on iOS.
@ChatGPTapp
If you need AI to do a search for you in the real world, ds4-agent is basically SOTA, because it can access the web sites without any limitations given that it uses your local Chrome browser (no, not in headless mode, that's the trick...), and DeepSeek v4 is great at search.
@antirez
One of the coolest in person demos at #wwdc26 was @lmstudio running massive local models on 4 daisy chained Mac Studios with a total of 2tb memory Then they pulled out an iPhone and chatted with those models remotely over a secure connection I need to find a use for this
@stalman
Claude Fable 5 has been out for a couple of days. Some projects people have already built with it:
@claudeai
Less than 24 hours ago, Anthropic dropped Claude Fable 5. Minds are blown. And people are already coming up with wild use cases. 10 examples:
@minchoi
i think the models are great and do amazing things day to day and then i go to use them as a brainstorm partner for creative work and they are just horrible, no amount of steering gets them to be even slightly better makes you think about what these things actually are
@thdxr
Claude Fable 5 is our first generally available Mythos-class model. It ships with new safety classifiers that may flag certain prompts in dual-use domains like cyber and bio. We've added fallbacks: a refused request retries on Claude Opus 4.8 instead of dead-ending.
@ClaudeDevs
Designing loops with Fable 5
@RLanceMartin
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward (imo of the same order as Claude 4.5 was in November), peaking especially for long problem-solving sessions on very difficult problems. You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!). The model still has quirks that people will run into and the safeguards are configured to be a little too trigger happy for launch, which can hopefully be tuned over time. I feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing substantially. You can ask for anything - explainers, visualizers, dashboards, bespoke single-use apps (e.g. a full wandb that is hyper-specific just for your project), you can 10X your test suite, auto-optimize code, run giant research projects with custom HTML for the results, anything! "Free your mind" (Matrix ref). Really looking forward to all the things people build!
@karpathy
Fable 5 is now available in Claude Code and Cowork Fable is the best model I have used for coding, by a wide margin. It is a big step up, enabling less prompts and steers, more efficient token use, better code quality, better tool use, more intelligent self-verification, longer running sessions, and higher trust & autonomy. Happy coding!
@bcherny
Introducing text-to-lottie: an open source skill and harness for generating production ready Lottie animations with codex/claude code. $ npx skills add diffusionstudio/lottie Prompts guide and repo in the comments.
@konstipaulus
For making agentic videos - I did a deep dive into @HyperFrames_ vs. @Remotion using @slashlast30days. Of course I had to make a video to go with the article! Take a look at both products. Which are you using?
@mvanhorn
i'm arguably one of the most prolific builders you'll ever meet. and i do it on a single 5x claude plan. 🫣
@Shpigford
This is wild. OpenAI just dropped Codex Sites. Now anyone can give it a plan, dashboard, launch doc or idea, and turn it into an interactive app with a URL. 5 wild examples:
@minchoi
Had Claude Code build a snake game where the snake becomes aware it is in the game and then... stuff happens. Some impressive creative decisions by the AI (& also some very AI ones), I just gave a first prompt and some feedback on the game as it went. https://snake-awakening.netlify.app/
@emollick