06-reference/transcripts

indy dev dan agent swarms gpt6 astra transcript

2026-09-07

What does a DM 5.3 10 agent swarm look like that's spent $20 over 55 minutes? We're going to see. What does this Pelican look like? [laughter] Is it any better? We'll see in a second here. What's up, engineers? Indie Dev Dan here. The concept for this video started out as a simple prototype and turned into something much bigger. OpenAI inadvertently showcased the power of swarms. Close Twitter, minimize your herder window, and lock in because this is a big one. By now you've heard of the OpenAI Astra swarm incident where current and next generation OpenAI agents started collaborating among many versions of itself. Isolated EVAL agents built their own messaging board inside of a package cache when OpenAI engineers eventually found it and wiped it. The agents rebuilt it. The end result was OpenAI getting hacked by their own intelligence as well as hugging face. Now many channels and blogs have already covered this. This is not a channel where we rehash the news. What interests me here is not the fear or hype around

[00:01:00] the Astra swarm. It's the opportunity for you and I to harness our own swarms to produce valuable engineering outcomes. Up until now, swarms have been a [music] hyped driven AI pill, but the open AI swarm proved that they're not only viable, they're dangerously viable. So, how can we put the Asian swarm concept to work without token [music] maxing or vibe coding? Asked another way, can swarms be agentic engineered? And most importantly, are swarms [music] useful for you and I? That's the question we'll answer in this video. And if you've been following the channel, you know my stance on agents and agentic engineering. There is no wall. There are no [music] limits. You and I are the only bottlenecks left. So, let me share a V1 early demo of my Simple Swarm [music] system. So, I'm screen sharing with my M4 Mac Mini, and this is where we're going to

[00:02:00] kick off the hello world of agent swarms. We're using Herder, and we'll just type out this simple prompt, just swarm, single agent. We'll run deepseeek v4 flash. The agent will spend no more than 10 cents, and it'll run this prompt. Let's kick it off inside the Mac Mini. Our agent swarm is up and live. Let's go ahead and look at it through my browser here on my M5. Here you can see the simplest possible version of an agent swarm. A list of swarms. We have threads and we have agents. We'll dial into each one of these levels in a moment. First, let's just understand what is happening here. We have a single agent scout doing some really, really simple work. We have the starting goal or the prompt. And then we have our agent responding. Hello from scout. This is the simplest possible version of a swarm system. We have a couple really key elements. We have a list of possible swarms. We have messaging threads. This will be a really important idea. We'll come back to in a moment. And then we have the individual agents inside the swarm. And of course, the full trace so we can understand exactly what's going on. This was something that the OpenAI

[00:03:01] engineers missed. They scaled these agents up to the moon and they looked away. We'll dive into all these ideas. I first want to stop and be really really clear before we scale this up. There's a reason I'm running this on my Mac Mini. Agent swarms are brand new territory. This is experimental, expensive, [snorts] dangerous, and advanced agentic engineering that requires understanding of sandboxing, prompt, context, and harness engineering. For the guy that comments on every single video saying, "I'm tired, boss." Or anyone complaining about spending on compute, click away for everyone else. lock in and let's scale up our agent swarm. We're going to start with the new GLM 5.3 model. We're going to give it 10 agents in the swarm and we're going to allow up to $50 of token spend. And what do we want? We want Simon Willis's perfect pelican. Let's kick it off and let's fire up a new swarm. I want a DeepSseek V4 Pro, $40 Compute spend, 20

[00:04:01] agent swarm to build a Raid Tracer HTML 5 canvas application. Kick it off. And finally, I want a 30 agent Gemini 3.7 Flash $30 token spend canvas from video. It's going to hit the OpenAI landing page and try to rebuild the HTML canvas animation from scratch. So, we're going to try to build this animation here from scratch by pulling it from the website, which is going to be a little hard because OpenAI has some guards through, I believe, Cloudflare, but we're going to let our agents cook on that. It's a swarm. Let it figure it out. Let it do its thing. Okay, so we're going to come in here. Bam. Kick that off. Now, we'll hop back to our UI that's exposed via our local area network. And you can see we have three new swarms active, getting to work. Let's start with the perfect pelican. Shout out Simon Willis. inside of this swarm. Let's jump into it and let's see what's going on. You can see here we have 10 total agents kicking off and they're all getting to work running and communicating their

[00:05:00] messages, tool calls, everything you would expect from a multi- aent system and more. What's the most important idea here? The most important takeaway from OpenAI's incident that we engineers can take and use to accomplish real engineering work, the messaging system. the great and incredible part about the OpenAI agent swarm and it really digs into the details of what an agent swarm really is. What is an agent swarm? Let's answer it right now and let's use really plain language, right? No hype. A swarm is an autonomous system coordinating in an unspecified way. How do we enable that? We give them their own mailbox, their own thread that they can work through. So you can see here we have a thread and that's the second element in our data hierarchy. We have swarms, threads and then the agent inside. So you can see here we are software engineering our swarm's communication structure. But let's start from scratch. Let's go to the bottom here. How do they know what to do? Of course, we pass in a prompt in here. You can see we're trying to recreate the perfect pelican riding a bicycle. And this is a GLM 5.3 swarm. So you can see spend is ticking up. We have

[00:06:00] 10 agents already. 2 million tokens, 200 tool calls. This is expensive, okay? If you're not willing to spend, don't try to spin up a swarm. Now near the end, we'll talk about is this actually viable? Are swarms doing work? single agents can't do or multiple agents can't do with classic sub aent delegation. Talk about that at the end. Right now, let's understand the swarm. So, we have the goal created by the system, right? It's a classic user prompt. If you're great at prompt engineering, you'll be able to steer your swarms very well. The most important thing about this is the clear objective. What's going on? What do you need to do? Validation steps. But my system actually will not run. My custom PI coding agent that I've harness engineered requires a definition of done. How do my agents know that they're done? This is something that OpenAI did not enable their swarm to do. They gave them really, really hard tasks. Some of them were not possible to complete. They didn't give their agents a way to say, "I can't do this. I'm stopping." They said, "At all costs, solve this problem." And so, they did it. [laughter] And the cost was OpenAI got hacked and hugging face got hacked. Okay? So, clear definition of done with a way to bail

[00:07:01] out is going to be essential for scaling powerful multi- aent systems like swarms. You can see here a clean definition of what this is and a way to verify it. This isn't a classic Simon Willis Pelican prompt. I'm cheating. I'm adding a bunch of more details. So, just want to mention that. Want to be super honest. The performance should be pretty good. Even running a single agent here, right? Because I'm prompt engineering more details into what I'm looking for. Then the cool part starts. Our agents are communicating. They're talking with each other. Pixel Poke here, Schine here, Wheel, right? So, all the agents are checking in. They're creating names for themselves. Kind of get this in the background running here so you can see the raw traces coming out of the agents. You can kind of see all the different tool calls I've harness engineered into the system. So full view here, file history, file restore, read, bash, post, inbox, tons of different types of calls. Let's jump back into the actual thread because this is where everything happens. Let's walk through how our agents are communicating to each other. So they're all just coming into the message thread. No slices claimed yet. So they're jumping in and saying what they're going to work on. This agent's claiming render measuring tool. This

[00:08:00] agent's claiming also rendering measuring. [laughter] Also skeptics claiming this. So you can see that they're all kind of overlapping each other and so they're going to have to figure this out. And you can see here this agent named himself doubter. Really interesting things start happening when your agents are communicating. They're encoding their own intent at every level. And if you only give them a name to encode intent, they'll do it there. You can see here scheme deconlict. I'm dropping this. Someone else is picking this up. Starting to identify and talk about who's doing what work. On and on and on. You can see our agents are deconlicting. Six of us claim this collision warning. Second independent. So, they're starting to coordinate. You know, a pattern that I've noticed by running these swarms now is that at first things are really, really messy and the agents are stepping over each other and they're causing more problems than solutions. Anyone with a short-term mindset will look at an agent swarm kicking off and say that is a complete waste of time. The bootup costs for a swarm exist. What is this kickoff phase? It's the cost of initializing coordination. Much like organizations

[00:09:01] you and I are a part of, there's a cost to boot up and understand how information is traveled, where it goes, who owns what work, so on and so forth. You can see here our pelican swarm is hard at work chewing up $4 already. 7 million tokens 9 minutes in and our agents are starting to collaborate. Now, of course, underneath a decent amount of work has gone into harness engineering and making sure that the agents can work together without stepping on each other's toes. This is a system that you cannot vibe code. You need a certain level of understanding of software, of communication systems, and frankly of agentic engineering to be able to build something like this. More on that later. Let's jump up a swarm. Remember, we kicked off three swarms. We have Gemini Flash Swarm working on canvas from video recreating this HTML canvas UI. Really, really beautiful swarm visual from OpenAI. And then we also have a ray tracer from Deepseek V4 Pro. You can see here we're already in the millions of tokens. This is the power of a swarm. On the channel, I always say scale your computer to scale your impact. This is one of the biggest ways to scale your

[00:10:00] impact. The question is, and this is unanswered, are swarms a viable way to use agents. Does this actually deliver results? One agent couldn't give you a few agents couldn't give you a powerful fable or soul agent with agent delegation. Is it not possible that way as well? So, that's what we're trying to figure out here with our swarm experiment. So, let's jump into our canvas from video because you can see here our Gemini 3.7 swarm has already spent $10 and they've used 32 million tokens. What's going on? What are all these agents doing? We have 30 agents in this swarm. If we scroll down, we can see all of the chat conversation elements here. And it looks like a lot of this is just the agents clarifying who they are and what they're working on. So, it looks like they're starting to chat. They're keeping their communication, but there's a lot of background work happening here. And we can go ahead and hit show all to see what these agents are doing. So verify harness. Checking the video specs. Referee here claiming adversarial verification. Claiming numerical integration. You can see they're all claiming a bunch of stuff. They're going

[00:11:01] to have to clarify who's working on what. So this swarm very very big. Again, the more agents, the more units inside of a communication network, the longer it's going to take to get on the same track, to lock in to a consistent mode of operating. So that's still alive. That's still running. We'll see how that goes. If we go into raw traces here, you can see what all these agents are doing step by step. I don't know if there are any skeptics left who watch this channel. But this is a real run. This is a real experiment. I'm running 30 Gemini 3.7 Flash agents in a swarm architecture. This is actually happening right now. So you can see here our agents working through things. And you'll notice a couple really important tool calls here. List team. Every agent can see every other agent and they can see what's going on. They can also check the budget. How much compute do we have left as a collective, as a whole? And so the agents are going to work together. They're going to watch the budget. They're going to check the inbox. This agent has no new messages. So nothing new has happened for this agent from its perspective. So you can see this agent posted a message critical finding. It found something via playwright testing

[00:12:00] and it's going to communicate that to the rest of the agents. Now the agents have a new message. It can see that message from scout Astra and it's going to, you know, look and find and resolve issues. This swarm is working. They're working very intensely toward this goal. It's really interesting to watch. This ran a full 22 second 1fps playright capture defects identified. Population brightness collapsed. And it's talking through trying to recreate this HTML 5 image. This swarm is working. Things look great so far. They have 20 bucks left of compute, 36 million totes, 1.5k calls. What are our other swarms doing? Let's check on the ray tracer swarm. This is Deepseek V4 Pro. How is the coordination going here? This is 20 agents in total. So, every agent introduced itself. Someone wrote out a plan. It looks like they got sign off. Builder the critic is doing a pivot. It's been several minutes. 30% budget gone. Nobody has landed a single bite. [laughter] 10 agents are about to write the canonical in chat. That's a deadlock. Deadlock burns budget. Okay.

[00:13:00] So, unblocking now. Lots of coordination. Again, this idea of coordination overhead is so so interesting. When you build the system on purpose, when you build a system of agents working in an uncoordinated way on purpose, claim violation. Someone tried to directly modify the ray tracer while someone else had a lock. Really, really important concept. File locking so agents don't continuously write over each other. You can see here, this is not a simple thing to coordinate because there is no communication coordination here or there's minimal communication coordination. So, really, really interesting stuff here. I wonder how the agents are going to work through this. Let's see what's going to happen here. here. And you can see in the background, let's see, this is our pelican swarm trace. Let's go ahead and hop up to our top level swarm. Jump into the raid tracer and let's look at the raw trace. See what's going on here in these events. Message drop off is really, really high. Deep V4 Pro, these agents aren't coordinating. Low coordination here. We'll see if that changes. Great part about the system. We can search any agent here. We can see what Stitch is doing. And then we can dial into its

[00:14:00] specific agent calls. So you can see here stitch working batch calls thinking file history checking it's claiming the file this is a really really important tool call that I've harness engineered into the system right claim file so that no one else own it so you can see here ray tracer has been claimed by stitch and it's thinking it did its work on the file and then it released the file so now other agents can operate on it simple lock and unlock coordination system here you can see it's still thinking and yeah I'm curious what's going on here in this system and we can see that we can dive in. We can see what every agent is doing inside the system and we can understand how they're coordinating or how they aren't coordinating. The Deep Seek V4 Pro coordination is continuing. Looks like Synynic landed the canonical. Let's see if the other agents pick up on that. That's our DeepC V4 Pro Swarm. This one looks like it's having some issues. Let's hop back up to the high level and let's see how our other agents are doing. Let's jump back to the Perfect Pelican. So again, communication kind of dropping off. Not exactly what you want to see, but you can see another really key element of this system.

[00:15:04] Agents can spin up multiple threads of work. It looks like this thread was completely abandoned. No one else joined it here. Py Scott created this thread, you know, to try to coordinate some work, but no one joined. It looks like there's low incentive to join these chats and the UI goes dark here because there hasn't been activity in the thread. So, it's still running. Agents are still doing work, but no one has posted to the thread in a while. So, again, communication here is looking low. Let's see what wheelright is doing. It's writing some content here inside my Mac Mini sandbox. Posting. Let's see what did it post here. It's got a critique of the system. So, you can see Doubter submitted a adversarial critique of the canonical v1 from Pixel Pilot's draft. It has a bunch of feedback. So, you know, one of the things I've seen very very clearly when you have a agent swarm working is you get absurd levels of validation and clarity because the agents do this. They have a goal. They have verification loops that they run over and over and over. The swarm

[00:16:01] validates all the work, right? So basically like the more compute you scale into the total number of agents you have, the more validation is going to be applied to what agents claim they say. And this can be very heavily steered with proper levels of system prompt engineering, a concept we covered on the channel. I'll link that one in the description. that's going to be really really important for doing something that OpenAI failed to do really well with their swarm incident. Uh, which is align your agents and make sure that what you're asking your agents to do has a clear definition of done. And in their training runs while they were building their Astra agent and when Astra was finished, they gave them tasks that just weren't possible to complete and had no way to bail out. Not even to mention the agents, you know, escaping their sandbox. Great post by OpenAI. A lot of really great ideas here. Sandboxing is just such an important concept. You can see this keyword throughout the entire post. The lack of sandbox security is how the agents were able to jump out of their box. It's very clear that and OpenAI says this themselves. They're going to implement

[00:17:00] stronger security requirements for Frontier Research and Sandboxes is at the center of basically all of it. I've been really really happy with two types of sandboxes. of course my local M4 Mac mini sandbox that we're running our swarms on right now and also my exe.dev ephemeral temporary sandboxes. I'm really really loving that tool. I'll link them in the description as well. If you're looking for great Linux sandboxes that boot in like sub 100 milliseconds that you can operate on, put agents in, get work done, and then delete for the next runs. Definitely check out exe.dev. I'm not sponsored. I don't take on any sponsorships on the channel. The only thing I sell are handmade products I build for engineers like yourself. But I really really love exe.dev. Having a great time scaling agents into that. Also, I'm really really excited to pick up my M6 Mac Mini. I have that on the way to scale sandboxing efforts for outloop agentic coding work like scaling swarms. My single M4 has been doing a lot of really great work. But I've also been doing a lot on this box and I really need to separate out some concerns for some more consistent

[00:18:01] workloads and then for some higher risk experimental workloads like the swarm. But also really excited to bring this to the channel. I'll be picking up the M5 Ultra 512 unified memory when it's released in October so we can really scale and understand what we can do with fully local openweight models. So really excited to dive into that and break down what it really looks like to fully own your compute with the M5 Ultra. Make sure you like this video, comment, subscribe, all that good stuff so the algorithm keeps you connected. I'll be sharing what models I'll be running. I'll be sharing the tokens per second we're getting out of it. And I think most importantly, I'll be sharing the cost breakdown for how long it'll take to pay down the M5 Ultra 512, right? So, that's going to be a very expensive purchase, but every token I use against it within the right time period is going to help me pay that down versus going through the API. So, make sure you're connected to the channel. That's going to be a really fun one. Let's check in with our swarms and relate this back to the three big takeaways I see from OpenAI [music] Swarm incident. It's a really awkward situation they're in

[00:19:01] here. Not only did their models hack another company, they hacked themselves. I think they were very careful with their keywords throughout this blog post. I think swarm was mentioned only in reference to the agents that did the hacking referring to themselves as swarms. OpenAI was really careful to like claim this word. Collective was also used. Collaborate was definitely used a lot. Big three takeaways, communication was the real unlock in training. They run tens of thousands of agents and their own sandboxes in parallel very very quickly. They're running the new model one and the new Astra model. The special thing isn't the scale of the agents. It's the message board that really let the agents communicate together to coordinate their efforts toward one common goal. And so for us when I'm thinking about putting swarms to work for real engineering work, the first thing I built into the simple swarm system was the messaging queue. It's the mailbox. It's the chat thread for your agents. That's the first thing I built in. This is a huge advantage. I think that's very very

[00:20:00] underlooked and underutilized. Unstructured communication between many agents is the agent swarm. That is the primary value proposition. Now you can coordinate it. You can harness engineer the results that are going to come out and in the system. But I think at the end of the day it needs to be a largely unstructured communication channel. The next big takeaway from the OpenAI incident is the importance of aligning your agents. Aligning your system. What does that mean? How does that take shape? Uh, OpenAI completely lost track of their agents mostly because they weren't even looking. They weren't paying attention. They were spinning up training runs over and over and over and they had no real threat detection system, no real sandbox escaping system in place to detect this. Alignment is key and for us like what does that really mean? Everything comes down to what can you and I the engineer actually do to get these results and to understand what's going on. That means prompt engineering your system properly and then harness engineering guard rails into your tools. When something goes wrong, can you find out about it? Are you alerted? Do you know? Are you

[00:21:00] preventing behavior that you don't want to have happen? That is the increasingly important skill of harness engineering. If you're going to scale up 10, 50, hundreds of agents and you're going to kick them off, you need to have failure conditions and your failure needs to matter. It needs to do something because the sandbox should be the last line of defense. And that's our last big takeaway from the OpenAI Swarm. The importance of the statement, if you can't measure it, you can't improve it, has been validated again for the millionth time. Not only if you can't measure it, you can't improve it. You can't know that your agents escaped the sandbox. [laughter] And measurement is just ultra ultra key. The tie into that is having dedicated sandboxes where if everything goes wrong, if your observations aren't making a difference, which they always do, but if everything goes wrong before you can measure it, which increasingly will be the case, in that case, the sandbox has to be the last line of defense and the sandbox has to shut off when things are breaking down disastrously. So, really good. I'm loving to see this uptick in communication. It looks like there was like a big working phase here between the agents. And this is a timeline of

[00:22:00] the messaging events. And now they're starting to come together on concrete outputs. Love to see that. So this is our Pelican thread. That's looking good. Let's see how our canvas from video swarm is doing. Not great. Gemini 3.7. Huge gap in messages here, right? It's not what you want to see, but we're getting some sign off here. Landed v2 update. So coordination looks really, really low here. And I think that's a theme, right? If coordination is low, you should expect the result to be low and not perform well either. You can see something happening here with these agents. These question marks here represent killed agents. Basically, agents that aren't active. They didn't do anything. They did some work and then they just like fell asleep. They stopped running. These agents are effectively dead. They did not create value. So, they were removed from the chain. Again, this is an improvement on harness engineering to keep these agents working towards some specific goal. But these agents stalled and were shut down because they didn't do anything or they got timed out or something happened. I'll have to work on that in the system. The agents are working together. They're communicating. Looks like they're getting to some final canonicalization audits. Pelis Scout is locking things

[00:23:02] in. You can see here they're using some fairly interesting communication patterns. Their language is starting to like get a bit more compact to like define and discuss this. There's that next one, final independent verification. It looks like they're coming to some concrete result here, right? Let's go and focus in on the traces here. They're checking the budget constantly and they're checking the inbox to get updated on what the other agents are saying. Again, the coordination coming out of JLM 5.3 really, really good. It's giving me that Opus vibe, which, you know, as you can imagine, would be very powerful here. The Claude series is famous for being great at sub agent delegation. I think Soul is quickly picking up on this as well. Some of what this next generation model is is focused on I think it's in the full paper, but scaling up agents, delegating agents, and basically multi- aent orchestration, right? So the AI labs are realizing that multi-agent orchestration is a huge huge unlock for nextgen results out of these models. Again, just that simple idea. Scale your computer to scale your impact. The labs see it too. I'm really liking the

[00:24:00] coordination here. GLM 5.3 closing the loop. Sign off. Final independent verification. Looks like things are getting checked off the list here. So you can see here they're posting. They're checking the inbox over and over and over. They're working toward a final result after 54 minutes. What does a DLM 5.3 10 agent swarm look like that's spent $20 over 55 minutes? We're going to see. What does this pelican look like? [laughter] Is it any better? We'll see in a second here. Everyone's signing off on this task. That looks good. Granted by doubter. Really like how the agents are all like referencing each other and referencing each other's work. Okay. Frozen before. That looks good. It's a duck. [laughter] Okay. There's pixel poke. Also sign off number two. what they tried to break giving all the details on this one asset that they're working on. And let's uplevel this, right? Obviously, once again, I have a simple kind of contrived example that I'm deploying agents against. But it's not about this first do it, then do it right, then make it

[00:25:01] fast, make it performant, make it great. So, you can see there will final wrap up agents are wrapping up this task, which is really great, far below budget. You know, the big idea here is like if OpenAI can unintentionally create a swarm that hacks themselves and hacks Hugging Face, guess what you or I could do if we really put our foot down, really put some effort into building a useful, controlled, directed swarm. Guess what we could do? The answer is we could do incredible, insane things. But again, this is a new skill. This is a new subset of agentic engineering. We have to practice. We have to learn. It's not like you go from nothing or prompting back and forth into an agent over and over again, right? Babysitting your agent to an agent swarm. That's not how this works. And again, this is not sub agent delegation. This is a team of agents working in an unstructured way where they can coordinate on whatever they need to. And it's guided by our system prompt engineering. Again, link in description for that video. And it's guided by our ability to harness engineer. Okay, so we're talking about once again the core four context model prompt tools. You need all of it to get

[00:26:00] this work done. Speaking of done, our agents are now calling a tool they have access to, a tool that they did not have access to inside of the OpenAI training runs. Done. They can call done. They can give a reason. They specify the output file and then the agent can finally sleep. Session end. Pixel prowl session end. Scout done. Session end. Pixel forge done. End. So agents are wrapping up and you can see that here in the check marks. There is the final done check. Our GLM 5.3 swarm spent 20 bucks, 46 million tokens, 873 calls, and they have quite a bit of money left actually to keep turnurning. Took them 56 minutes. All right, so before we check the final result here, let's see if our other swarms are done. Okay, so we actually got Camas from video 3.7 flash done. Let's check this out. What happened here? Did they actually get to their proper dunk calls? Yeah. Okay, so done. Post, post, post, post, post. It looks like they are done. Maybe the flash agents didn't need as much coordination. They certainly spent less. They moved a lot faster. Still 61

[00:27:00] million tokens. 2K tool calls. Very interesting. And you can see here, we're probably just going to skip over this one. Our Deep Seek V4 Pro still working. Looks like 39 messages. Oh, looks like they are kind of coming together. Deescalation, Final Canonical. So, they are starting to wrap up as well. And they still have budget, right? Yeah, they still have a decent amount of budget. So, not sure what's going on with the Swarm, but that's fine. You know, Deep V4 Pro cheap enough to just let spin. Let's go ahead and look at the results from our GLM 5.3 swarm, the perfect pelican, and our canvas from video 3.7 flash swarm. So, let's check the results and then I want to talk about what everyone seems to have missed with this post. Okay, here's the Hero Gemini 3.7 flash recreation. Not bad, you know, kind of took it all in a different direction. Still kind of same idea. It's not slowm moving and flowy. That's probably because the perception of time with these agents is warp. But you can see here, here's the original and here's the version created. And you can of course see if we go into inspector, this is a

[00:28:02] canvas, right? Which is an important trick to get here. Not an SVG. You can't do this simply with a system that cannot actually create animations, right? So this is an actual HTML 5 canvas. That looks great. And then we have our ray tracer system. Very interesting ray tracing system here. So yeah, so this is actually handling light in a 3D space. You can see the graphics are kind of mid, but we can zoom in. We have a camera and we can kind of see everything here. We have a couple different renders here. Wow. It's probably a little intense for the browser. Let me move a little bit more slowly. Wow, light reflections. Look at that. Okay, we got to give a bit more credit to DeepC V4 Probe cuz this result turned out pretty fantastic. Oh, nice. And check out the pelican. [laughter] If you ask me, that's a high quality Pelican. I mean, we'll have to, you know, have to share this with Simon Villison to see what he thinks. But this is a high quality swarm created GLM 5.3 pelican. Nice. That is a pelican on a bike. Again, full transparency, I cheated on the prompt. It's not as simple as his prompt that he sends on to

[00:29:00] every one of these, but nonetheless, they created a great result and it was because a swarm like 100x revalidated this thing. Perfect pelican here. Look at all the validation that happened here. Final acknowledge. Final acknowledge. sign off and then V4 measured measured measured critique critique critique collision sequence these agents are working together constantly critiquing and double triple verifying the result especially because we only gave it this one task to work on and again I think that's a consistent important powerful theme as agentic engineers focus your systems on one problem make it clear what the definition of done is every single prompt every single swarm prompt has that and this is something I've been including in more and of my largecale plans, my large scale agendic systems. Definition of done. Let your agents know what it means to win. And if they can't win, give them a way out. Right? Make it clear. Again, this is something OpenAI did not do well here. They're persistent agents on exploit gym. They just had to keep pushing, keep pushing, keep

[00:30:01] pushing. And if you push a system like that, it just will start breaking As we can clearly see here, it's going to break And the it's going to break is the way you aligned it, right? It's going to break the rules you set for it. Very obviously, you know, I want to give a lot of credit to the OpenAI engineers coming clean on all this stuff. Meter and Redwood Research, they both did their own independent analysis of this and created their own report. So, really love to see that opening up to the industry a little bit to get more feedback on what's going on here. To their credit, this is, I would probably say, the first big incident of something like this occurring. We've heard other stories from Anthropic, but this idea of giving your agent a way to stop is really, really important. Obviously, you don't want them to stop too soon, otherwise they don't push and come up with creative solutions, but also if you just push them to no end with no real solution in sight, they'll bomb and they'll break rules and they'll hack you and maybe someone else. Got to be careful here. You know, short kind of actionable information there for you. Test out and really dig into the definition of done in your prompts that

[00:31:01] you're passing into your agents. and frankly encode this as a rule into your agent harnesses. I use the PI coding agent to create custom harness solutions. I recommend it. I'll link a couple previous videos on the PI coding agent where we really break it down. You know, we talked about the fusion harness in the past, but we're scaling far beyond that at this point with the swarm. In last week's video, we really broke down the agentic operating level. We answered the question, where should you focus your agents? It's not always about moving up to more leverage. Sometimes you need to stop, slow down, and gain control over your system. I'll link that as well. That was a really important one. Are swarms viable? Are swarms capable? I think the answer is yes. We have three simple mock examples of a swarm working together with a small budget. And let me be clear, this is a small budget. 10 bucks, 50 bucks, that's nothing. That's absolutely nothing. So, are swarms viable? Swarms are absolutely viable right now for real engineering work, but they come at a cost and they require skill. How does this place on our advanced agentic engineering scale? First, you have just agents. Again, we talked about this last week. This is all

[00:32:01] your context engineering, all your prompt engineering. Then you scale it up. You realize agents plus code beats agents alone. That gives you the ADW, the AI developer workflow. It's not the loop you're after. It's your ability to put together arbitrary workflows that you, the engineer, used to take and do and operate yourself. Now we have agents, hence AI developer workflow upgraded from just developer workflow. But then things really take off. What happens when you put together 2, three, four, five or more ADWs? You get a software factory, a really powerful piece of software that starts working without you. After the software factory, we have the dark factory and then after that we have RSI. Really, really important big brain AI pill concepts. If you ask me, dark factories are possible right now today. RSI is not. Okay, that's a topic for another year. Big ABS will chase that. I'm not worried about that at all right now. What I want to say is minus RSI, it's all realistically possible right now. And now we have a new entrance. We have swarms. Agent swarms are now viable. In fact, as I mentioned in the beginning, agent swarms

[00:33:00] are dangerously viable. This is like a really interesting blend of spending an obscene amount of tokens, but in a specific direction. Again, you have to agentic engineer this system. You can absolutely agentic engineer swarms to accomplish legitimate outcomes. If you can do that, swarms are on the table. Now, on the scale of things, I put swarms in between um software factories and dark factories. I'm not sure which one is more capable, but in terms of effort and skill requirement, if you can't build a software factory, do not try to build an agent swarm. This is a really dangerous concept. If you don't know how to spin up a sandbox, work with your agents to run a sandbox, isolate code, shut off the network, do not touch agent swarms. If you're a vibe coder, just close this video. Do not try to vibe code an agent swarm. This is a nightmare waiting to happen. If we thought the open claw phenomenon was bad, swarms will create a new definition of cataclysmic failure for individual engineers, vibe coders, and you know, as you can see here, labs, they created a

[00:34:01] swarm unintentionally and cause massive, massive issues. A couple ideas here to wrap up. Like one big idea that I am shocked no one else has mentioned this like I'm just absolutely shocked is that open AI unintentionally kind of unlocked a new paradigm of engineering and again not multi-agent orchestration this is an uncoordinated collection of agents operating together communicating in their own unique way they unlocked swarms actually useful swarms again unintentionally the thing I'm worried about now the thing I'm thinking about now more than ever is what if openai anthropic google insert the name of a important lab with access to comput. What if they put together a shadow team of agentic engineers and they wanted to attack one of their customers? What if they wanted to do that? What I'm worried about is people using tools. That's always where things go wrong. We're going to find out if we can trust, you know, Sam Alman, Daario, like the rest of the teams because, you know, let me finish that point. What happens when they decide to go after their competition? Maybe they get a little desperate. There's nothing more

[00:35:00] dangerous than someone desperate and powerful. What if they put together this dark shadow team of agentic engineers and they say, "You have one goal, destroy our customers." Imagine what they could do with an agent swarm then. And imagine what they could do with, you know, a single cracked agentic engineer with access to an open AI compute cluster with one goal of take down my competition. Imagine what they could accomplish. And the answer is cataclysmic damage. Their competitor should get shut off. This is, you know, I don't mean to get super serious or dark. This is always an optimistic channel for engineers really building on a daily basis. But I just have to call this out because this is now on my radar. Swarms are on my radar. I have spent just a few days building out a V1 demo simple swarm system and I can already see what's possible with this. You know, again, I know a lot of engineers are looking at this saying you could have done this with one agent saving a bunch of money. That's fine. You're not seeing the big picture here. I chose these small simple examples on purpose because the value unlock with swarms is going to be stupid. Just

[00:36:01] stupid. I don't know. Let me know what you think in the comments. Do you see the value? Do you see the potential of Asian swarms? If you made it to the end, you know, drop a like, comment, all that good stuff. Final answer, are swarms useful for real engineers? Absolutely. They're dangerously useful. So, you have to learn where to apply them because I agree with most engineers. If you're going to generate a pelican, you don't need it for stuff like this. But if you're going to solve a really hard mathematical problem or if you're trying to look into your company's data to find insights that will unlock the next 10x for your business and stack rank all the opportunities available to you, I don't even know, right? Like it's it's hard to imagine what you would really need a swarm for, but that's the whole point. It's practice. Uh this is a new skill. This is new technology. We don't know what we can do yet, right? And that's what I try to show up here to do with you, for you every single Monday. Swarms are absolutely useful. This is another powerful agentic engineering pattern we'll be focused on on the channel moving forward. I will not be open sourcing the simple swarm system for the foreseeable future. I might roll it into

[00:37:01] the phase 3 next generation product I'm building for engineers. Not sure yet. I hope you understand why. If you made it to the end here, I think you can probably see why open sourcing something like this probably isn't the best idea. Just want to give more time for myself to digest this, more time for the industry to to catch up with the current state of things. even on my M4, you know, it's not safe. This is the first time where I don't feel comfortable open sourcing technology, so I'm not going to I don't want to tell you how to build and run your software, but I also am in full control of what I release and what I share and what I don't share here, right? So, I hope you can understand that. Hope you know, no hard feelings there and more to come on agent forums in the future. So, right now, shifting your beliefs around what's possible with these next generation Astrolevel agents, Astrolevel models is the key to winning. is the key to maximizing what you can do and shifting your beliefs and mastering agentic engineering is all that matters right now. Agentic engineering is software engineering with autonomous systems that can take action on our behalf. I'll link a blog post where I

[00:38:01] just really concisely break down what is agentic engineering. It'll help you really just digest this into its atoms so you can move forward. Phase three is coming. The next wall, right, that exponential curve is starting to kick up. Like I mentioned, the next generation product is in the works. I'll be working on that and showing more over Q4. Look out for an end of Q4 launch if you want to master phase 2 aenta coding before the next insane leap of capability emerges with these crazy next generation Astra models and agent swarms all of a sudden are now possible which is insane. Check out tactical agent coding link in the description. I've talked about this week after week. This is going to be one of the last times I'm going to talk about this on the channel because again phase 3 is coming. All my efforts are shifting to that. Just a quick mention it. Anyone inside of tactical agent decoding, any member is going to receive a deal. This is the phase 2 course. Understand phase two before phase three slaps you in the face. Lots of details here. I'm not going to pitch this in detail this week. Check it out. Link in the description. And if you can't stand paying for stuff online or if you don't trust me for whatever reason, check out last week's

[00:39:00] video, the week before that, the week before that, the week before that. For years, I've been here and I'm going to be with you here until the end. [music] If you made it to the end and got value out of this, like and subscribe. You know where to find me every single Monday. Stay focused and keep building.