Three video models shipped audio in 72 hours
I went to check one video model this week and found three, all of which quietly stopped needing a separate audio pass.
FLUX 3 Video from Black Forest Labs went generally available on Aug 4. Clips up to 20 seconds, 720p or 1080p, native synced audio with dialogue in 14+ languages and actual lip-sync, multiple shots and camera angles inside one generation, keyframes, and continuation from four seconds of existing footage — video and audio. It also renders typography inside the scene, which is the thing that usually sends you back to After Effects.
Two days later Seedance 2.5 went live in Luma: 30-second single-pass generation with no stitching seams, up to 50 reference inputs per generation (it was 12), synced audio in the same pass, and targeted re-render of one object without regenerating the shot. ByteDance claims about 20% better prompt adherence.
And MiniMax H3 — open weights, 2,939 likes — took over the Hugging Face trending board. The base model landed Jul 28, but the derivative wave is what's interesting: Turbo LoRAs, quantizations and ComfyUI ports were still shipping on Aug 5, 6 and 7. It does 4–15 second video with native stereo audio up to 2K. It also went live in Luma Agents on Aug 6.
Separate voiceover and soundtrack pipelines are basically legacy tooling now. That happened in a week.
Draft Mode is the part that matters
Everything above is impressive and most of it doesn't change what you do on a Tuesday. This one does.
FLUX 3 Video has a Draft Mode that returns a cheap, fast preview and then renders that same preview to full quality — same subject, same composition, same motion. So you can generate 12 hooks for pennies, pick one, and render it properly. That's the difference between one-shot AI art and something you can actually iterate on.
Draft Mode plus keyframes plus multi-shot plus native dialogue is functionally a VSL engine. Nobody has wrapped it in a marketer's workflow yet — hook variants, offer inserts, aspect-ratio fanout, push to the ad account. That wrapper is maybe a weekend against the BFL API, and it's worth more than the API access it resells.
One practical note: Runway is giving unlimited Seedance 2.5 on new Max plans until August 14. If you want to test the two against each other, that's the cheap way and the clock is running.
Hedra turned its models into an API
Hedra for Developers shipped Aug 4 — unified API, native SDKs, a CLI, and MCP support. So a Claude or Cursor session can drive it directly.
The interesting bit is composition. An agent can generate reference frames through routed third-party models — Nano Banana, GPT Image — then call Hedra's own audio and video models in the same session, one contract, one bill.
Talking-head content has been a click-through-the-UI job. Now it's programmable. That's the difference between "we can make a video" and "we can make 200 personalized ones." Pricing isn't stated in the launch post; the API is metered.
Everyone started building for agents, not people
This is the actual story of the week and it wasn't one launch, it was four independent teams making the same bet.
- Nitro 4.0 (#3 on Product Hunt Aug 7, 215
upvotes) is a 15-year-old human translation service — 80+ languages, 63% of orders done in under two hours. It just became the first one to accept agent-initiated payments. An agent hits the API unauthenticated, gets an HTTP 402 with a challenge, fetches a one-shot credential, retries. One credential per request, so there's no stored API key to leak. About $0.10/word, no subscription. This is the most under-rated item on the whole list — an incumbent service business made itself agent-purchasable and nobody else in its category has.
- Aveiro exposes roughly 56 MCP tools so ChatGPT,
Claude or Cursor can publish sites, articles and newsletters directly. Starter is $12/mo, Professional $99/mo. It's a publishing target for agents, not a tool with an AI button on it.
- UCP Radar scores a Google Merchant Center feed out
of 100 and puts 60 of those points on "Agent-Ready" fields — material, product details, age group, highlights, FAQ. The ones AI shopping agents read and almost every feed leaves blank. $39/mo for 800 products up to $599/mo for 35,000. The rubric is the product; the rewriting is commodity. Maker is anonymous, so weigh the traction claims accordingly.
- Cloudflare OS — #1 on
Product Hunt Aug 6 with 461 upvotes — is Cloudflare open-sourcing (Apache 2.0) the agentic workspace it runs its own company on.
You don't have to deploy Cloudflare OS to get something out of it. Two design decisions in there are worth reading the code for: agents are never handed credentials directly, and there's an audit record of everything an agent reads. Both are free to steal.
The developer tools with real free tiers
Superlog Responder (#3 Aug 6, 294 upvotes, 41 comments) plugs into the Sentry or Datadog Slack channel you already have. On every alert it investigates with logs, traces, recent deploys and past threads, groups duplicates into one incident, then replies in-thread with a root cause and a mergeable PR. It's Apache 2.0, about 1,240 stars, and the free tier is 1M spans, 5M logs, 30-day retention and 50 investigations a month before it costs anything. Team of two, YC Spring 2026. Watch the domain — superlog.dev is an unrelated Git GUI from 2022.
Coldtea.ai was #1 on Aug 7 with 405 upvotes and 69 comments, which is the best comment-to-upvote ratio on either day and usually means people actually showed up. It bundles three things you'd otherwise buy separately: a terminal running your CLI coding agents in parallel with shared context, visual QA agents that drive the real app against preview URLs on every PR, and production monitoring that explains breakage in plain language. Free forever for solo on macOS with 2,000 credits, roughly 200 test runs. Pro is $20/user/mo with the first seat free. Built by Ohans Emmanuel over about five months.
The visual-QA-on-every-PR part is the hard one and the one everybody skips.
Model updates you can skim
- OpenAI, Aug 6. GPT-5.6 Sol got a slider for how much thought goes into a
response. Free users move to GPT-5.6 Luna as default with unlimited text chats and a Think button. A user-facing effort dial is table stakes now.
- Meta, Aug 5. Muse Code, a terminal coding agent with persistent background
sub-agents and an append-only event log that makes runs replay-exact. Meta demoed GPU kernel optimization over 1,000+ tool calls across 24 hours. Pricing is about $1.25/M input and $4.25/M output — or $0.10/$0.20 if you let Meta train on your code. It hit #4 on Product Hunt with 240 upvotes and three comments, which tells you how much developers cared.
- ElevenLabs, Aug 3. Big changelog, all infrastructure. Two breaking changes
to note if you're on it: CharacterAge middle-aged became middle_aged, and audio_filter is gone from the Agents TTS schemas.
- Suno, Aug 7. Voices on iOS and Android — record once, reuse on any song,
free plan included.
Nothing genuinely new from Midjourney, Ideogram, Recraft, Runway's own models, HeyGen or Synthesia. Kling's release notes have been unreachable since the klingai.com to kling.ai move, so I don't know either way.
Two things to be careful with
Soloop was #2 on Aug 7 with 341 upvotes and the positioning is genuinely different — an AI CEO, CTO, CMO and analyst working one company goal, with anything needing judgment routed back to you instead of auto-executing. But there's no public pricing page, and the landing page testimonials repeat in a loop and read as placeholder copy. Interesting idea, unproven product.
Framer AI Agents shows up on both days' boards with 883 upvotes. It's a promoted placement, not a launch — Framer 3.0 shipped June 16.
So yeah
Four teams, one week, no coordination, same bet: the customer is an agent. A translation shop taking HTTP 402 payments, a publisher shipping 56 MCP tools, a feed scorer putting 60% of its grade on fields only a machine reads, and Cloudflare handing out the workspace it runs internally.
If that's right, then the thing worth auditing isn't your site's Lighthouse score. It's what an agent sees when it reads your pricing page and can't find a number. I don't know what that audit looks like yet. Probably a crawl, a rubric, and a list of missing JSON-LD. Which is a rabbit hole for another week.