Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
May 06, 2026
·
Raleigh
Stateful agents with open-strix
Explore stateful agents with open-strix, a minimalistic harness using cybernetic principles for a solid, extensible core, covering memory architecture and viable systems.
Overview
open-strix is a minimalistic open source stateful agent harness that leans on the Unix principle and uses cybernetics principles to build a tiny, solid, extensible core.
Video
Transcript
Generated 3 months ago
Summary
Generating a talk summary...
View full transcript
Speaker 0: Empower is but I'll let Nick our friend Nick explain Empower here. I don't
Speaker 1: think I need the mic. Can y'all hear me? Yeah. Good. Thank you, Danny.
Speaker 1: So I'm Nick Allen, and I am the executive director for Taylor nonprofit organization called NPower North Carolina. We're part of a a national network of of of NPower in 14 markets across the country. And what we focus on on the surface, right, is that we provide access to free training for young people in the military connected. I means, veterans, veteran spouses, active duty spouses Way are looking to launch careers in technology. Right?
Speaker 1: I said on the surface, we we're a workforce development organization because what we really do is that we help our our our students build community. I? I heard Danny talk about the powerful community and connection in this Loop. And y'all y'all understand the the the how important high value professional networks are to achieving your goals. I?
Speaker 1: So for our folks who may not be on traditional 2 year and 4 year degree tracks, how do they find these networks and connections and community to help them grow their professional careers and their understanding and their knowledge? I? And so it's through partnerships, with folks like Danny and other community partners that we're able to provide and bring folks into our space and invite our students in, and they start to build their own connections and knowledge and I professional networks to help them launch and sustain a career in technology.
Speaker 2: Right?
Speaker 1: And And so we're excited about the opportunity and the connection. We even talked about our version of, like, what does it look like to have a Tinkerers community when you're first starting out? And that's what we're looking to build. And and so if you have any questions, look us up. Please connect.
Speaker 1: Danny has my contact information, but we love to to bring people into our space to share ideas Two help our entire community, leverage all that Triangle has to offer
Speaker 0: And if you're looking for a good volunteer opportunity, by all means, feel free to volunteer with Empower. Yeah. We've got I my friend, David, Eric here. We did a the Center for practical I. We did a beginner AI course at the Empower campus– there.
Speaker 0: So we were showing folks who are not AI fluent like we are, is how to understand what an LLM is and break that fear and understand, you know, hey. You don't need to be scared of AI. Here's the things you can start doing today to use AI. And so, definitely a great place to volunteer because, the work they're doing is amazing. And, you know, we're privileged because we're on the other side of the AI, you know, field where we're we we know we're entrenched in it, but it's great that we can give back to those who are are learning and and upcoming.
Speaker 0: So our future tinkerers are gonna come out of groups like the students had in power.
Speaker 3: Thank you, Danny.
Speaker 1: And and just for logistic reason, we're located in RTP. Right? We're right there on the Frontier Campus in Richmond. The box yard? Yeah.
Speaker 1: Yeah. Yeah. Yeah. Yeah. Right?
Speaker 1: We're right across the street in the 800 Building on the 4th Floor. And the idea around that was our folk we could've went to Southeast New Raleigh, East Durham, right, where a lot of our folks typically come from. But we wanted to have Make it work. An environment that mimicked the future work environment that we've prepared people for. I?
Speaker 1: So we're in RTP, and so I they can start to build that confidence and sense of belonging that we know is important to to be successful in a career. So come check us out. RTP, come get a video at full steam, and learn about that file.
Speaker 0: Plus, they have the coolest TV. And I we got to be the guinea pigs that day. Yeah. So when you ever see CNN and they're touching the screen and moving in and doing this, they have 1 of those there. Yeah.
Speaker 0: 2
Speaker 2: of them.
Speaker 4: 2 of them.
Speaker 0: 2 of them. So, it's it's really cool. So Yeah. Yeah. Thank you, Nick.
Speaker 0: Appreciate y'all. Thank y'all. Yeah.
Speaker 5: So once we get
Speaker 0: more details on that, we'll go ahead and and let you know what's going on for that. But June 6, that one's gonna be a little bit Taylor, so we're gonna have to cap our RSVPs. Oh my god. So just talking about our leadership here. For some of you, you might remember Paula Ramos.
Speaker 0: Unfortunately, Paula, with her commitments to NVIDIA, is not able to help out as much with tinkerers anymore. She's still part of this community and still be helping out and and that. But, right now, I've kind of stepped up to lead this. And, Tanya, can you stand up? If you see anything on LinkedIn, that's her.
Speaker 0: So she's been great at helping us with the social media, making sure everything is is there. And her husband is as well. Big part of the community. You may have seen him at all things AI. I know he was there as well.
Speaker 0: So we appreciate coming all the way from Greensboro and helping and and supporting our community. And, of course, Ron, who everyone knows, do we tell him, Ron?
Speaker 2: We'll we'll do it at the end.
Speaker 0: Alright. But, but yeah. So those are our Linkedin, so you can I I'm sure I probably connected with, like, 70% of everyone already here on LinkedIn? But if I'm not, feel free to add me, New, and Ron. And, we'll be more than happy to keep talking because you probably connected me on LinkedIn or you've seen me because I'm, like, I'm everywhere.
Speaker 0: So you've probably seen us, so feel free to connect with us afterwards. I.
Speaker 2: So I
Speaker 0: wanna thank our sponsor, Zed, sponsored this room for us to be here. So we thank them for that. They were, kind enough to, come in because, honestly, it wasn't cheap, but they were able to to help us with this. Raleigh founded, of course, a big supporter and (PostHog’s, with the generous donation of of pizza. Because Mark Kiko always says, you know, community, you make great friends, but pizza's expensive.
Speaker 0: And so, we we appreciate Posthog, for for giving us that wonderful donation of, the I of New York and all the pizza, the wings, the soda, and and everything. So we appreciate all of our sponsors.
Speaker 2: Posthog has a really cool website too. You definitely should check that out.
Speaker 0: Yeah. Alright. So to the meat of the matter, our presenters. Two for those presenters we haven't talked to, because we have so many, we're gonna limit the talk today. So we're gonna let you know when you hit our 2 minute warning.
Speaker 0: So don't be surprised when you start seeing signals to get ready to to cut there. By the way, real quick, how many of you it's your first time at a tinkerer's meeting? Wow. Wow. Okay.
Speaker 0: First off, thank you. We appreciate you coming out. So for those of you who've never been to 1 in 4, we basically have our demos. People will show you what they're building. Sometimes it doesn't work.
Speaker 0: I, last time, we literally had poor Mark Serkin who
Speaker 2: could not
Speaker 0: get his demo to run, and that's what happens. For anyone who's wondering, should I demo? It's not polished. It's not ready for prime time. Don't worry about that.
Speaker 0: If you've got something you're working on, an AI that you wanna show off, by all means, feel free to come, submit a talk. It's kinda like Saturday night live. Saturday night live goes on the air, not because they're ready, but because it's 11:35 eastern time and they need to get on the air. Right? So it doesn't have to be perfect.
Speaker 0: We're not looking for perfection. We're a community that people are building and they're sharing, and you're gonna get feedback. So if you're building something and you wanna know, does this look good? Does this seem like something people would want? This is a place to come demo and and show off that and meet other developers.
Speaker 0: And if you get connected with our global we've got the Discord. We've got the message board. You'll connect with folks I. And all of these talks are posted on YouTube. And as a matter of fact, any Tinkerers DeepAgents across the world, those talks are posted on YouTube as well.
Speaker 0: So you can find the best of. So if you're doing something with voice agents, you can go to their YouTube and see what people are doing with voice DeepAgents, imaging, processing, multimodal, whatever it is, you'll find it there. So, normally, we give the best demo award. We are gonna give it today. We don't have the card with us in person.
Speaker 0: So whoever wins it today, we will have to make arrangements for you after the fact. But, we do have our award for the best demo, so keep that in mind as you're watching the demos. Afterwards, we're gonna vote, and you decide who will get our best demo award. So with that, we're gonna leave it to Ash who, Ash has been a supporter of ours in the very beginning. So we appreciate Ash is right away, he's like, yeah.
Speaker 0: I'll demo. Right away, he's always there to support and, you know, be a part of our community. So we appreciate Ash. So, Ash, the floor is yours.
Speaker 2: Do Two have H 10? Right?
Speaker 5: Yes. That's a small 1, I think. It's gonna be On this 1. Oops. It's the other side.
Speaker 0: Oh, and while he does that, if any of you take pictures afterwards, on the Tinkerers page, you can feel free to share those because we'd love to see your pictures as well. I've definitely took a bunch of pictures, especially Loop Nick, budging around with the VR headset. But please feel free to take Tinkerers, do that. We have our LinkedIn page, so feel free to tag and connect with our our our presenters and and share on LinkedIn and social media what we're doing. We'd love to we'd love to see our pictures from our community.
Speaker 0: So thank you.
Speaker 5: Alright. Hey. Thank you everybody for being here. I'm gonna be quick because I have only 10 minutes. So, my local setup isn't working for some reason, so I'm just gonna have to demo the live version.
Speaker 5: So it's basically we have, things at our, homes. Right? If you if you have ever, seen George Carlin, he has a very, funny bit about stuff. So this, app is basically about stuff that we all have at home. So when I'm looking for something, I I know I have it somewhere, but I can't find it.
Speaker 5: So I've been, trying to find a way to make it easier for myself to I things that I'm looking for that I I can't find. Right? Sometimes, I even end up going to Target or, you know, Home Depot and just buy another 1 because it's just a hassle to go through and find it in my boxes. So this is basically a multimodal RAG application. It's, sitting behind a very simple user interface.
Speaker 5: All you have to do is take a photo and upload it, and then you just forget about it. Right? You don't have to tag it. You don't have Two, make a list of things, in the boxes. So you're basically using a multimodal embedding model or a language model to, you know, do the work for you, basically.
Speaker 5: So the let me just, do a quick demo here. I was gonna run it locally, but it's not connecting. It's, it's an application built on Microsoft stack. It's a Taylor application, if you know about what that is. It's, using Azure, storage, SQL Server, that that stack, basically.
Speaker 5: Right? That's not important, though. I'll come back to that in a second. So, you can see here on the screen I'll blow it up a little bit. There are photos of boxes with a lot of stuff in them.
Speaker 5: Right? So let's say, I'm I'm looking for, this thing here, the Groot. Right? So I'm just gonna say Groot and it goes so while it's doing that so it basically lifts things in the order in which it's most I, to be in the photo. Right?
Speaker 5: Let's Way, so this you can see here, right here. Right? The same toy is in this picture, but it's showing very differently. It's laying on its side. It's barely visible, but it's still able to find it.
Speaker 5: So this is basically the functionality of the application. Right? So it's very simple. You can do things, for example, in this case let me see. So in this box, I have a Fitbit, and I've put a sticker on it with a label that says number 2.
Speaker 5: Right? So and this I can, obviously, I can I can enter the label myself, but you can also just, let me see if it's, New? I'm not gonna be able to show that. That's on my local. You can use AI to just, like, find the label and label it as well, like, if you have the labeled boxes.
Speaker 6: So
Speaker 5: the the main thing here is that, New are using multimodal embedding models and language models to do this. Right? And what I showed you is basically all you do. So if I'm doing add add let me see. Maybe this 1.
Speaker 5: I can also do things I, if I don't, specify, specifically that I'm looking for this 1 specific toy, the Groot. Right? I can also Way, I don't remember the name of, the toy, but it's a science fiction movie character toy. Right? So that's enough for it to, you know, find that and other you know, what whatever else that fits that description as well.
Speaker 5: Right? So it's obviously, using the, voice, natural language input to create those prompts. So for example, I just loaded this photo, and I'm just gonna say I red hat. So because I just uploaded it, I just wanted to see if it's, okay.
Speaker 2: No.
Speaker 5: Okay. So so it's it's working now.
Speaker 2: Way model is using?
Speaker 5: What model? It's, for some reason, it's not working like that. Yeah. Yeah. That happens.
Speaker 5: Demo guards. Yeah. Alright. So, it's using I can switch. It's sitting behind a router.
Speaker 5: It's not open router, but it's something that I built myself. So I can switch the providers. Embedding models, I've, experimented with OpenAI, of course, but also, Cohere, Jina, couple other ones as well. So, Jina and Cohere's multimodal embedding models are really, really good. So, in fact, in this demo right now, that's all I'm using.
Speaker 5: I'm not even using the language model.
Speaker 2: Can I ask, so with the multimodal embedding model, if you're not it's labeling the images as well as allowing you to do image search? Like, it knows that the red hat in there. You're not doing a separate labeling step?
Speaker 5: Yes. It just works just with embeddings. Yes. Yeah. So how much time do I have on?
Speaker 5: 3 minutes. 3 minutes. Okay. So so behind this very simple user interface, there's a lot of engineering effort that went into it. So it's an actually production ready application.
Speaker 5: It's deployed in Azure, accessible to everybody. In fact, if anybody wants to try it out, let me know. I can let you in at no charge. But it's when you're looking at creating a you know, if you're working on your own project, right, you want it to be scalable, resilient, you know, secure, all those things. Right?
Speaker 5: So this application covers all those areas also. So in fact, let me see. So this is an example of the techniques and patterns that have been used in this application. So for example, envelope encryption. Right?
Speaker 5: So I wanted the photos to be private. So if I'm uploading my photos, I want them to be private. So I encrypt them before they are saved. Right? So I'm not using just the encryption at rest provided by the cloud providers.
Speaker 5: Right? I'm encrypting these photos with your password. And I'm using the same mechanism that, Bitcoin uses to generate its wallet, keys, the, mnemonic keys, if you've ever used it. So it's the same mechanism that I'm using. So if you lose your password, there's no way I can decrypt your photos.
Speaker 5: Like, I cannot see your photos. They are encrypted and decrypted on the fly. So, like, when I'm do when I'm going through this u I, right, for example, if I so if I page Two this, so it's on Two, page 3. This is being decrypted on the fly.
Speaker 3: You got 1 more minute if you
Speaker 5: want. Sure.
Speaker 2: So,
Speaker 5: I. I can I'll take questions now. Thank you. Go ahead, please. That's a great question.
Speaker 5: So I actually I I did run into that very quickly. Like, I forget my password. Right? So, a lot of times I used to, not anymore, but I would just go to the website and say, change my password. I, I don't remember it.
Speaker 5: Every time I go there, I would change my password. Right? So so the problem with that is that you have to, like, now, handle that scenario because I've already encrypted 3,000 photos of yours. So how do I do that? Right?
Speaker 5: So I need to know the password to decrypt them, to Two to reencrypt them. Right? So that whole situation gets very messy. So the way you do that is by using something called envelope encryption. So I'm encrypting it with what's called a data encrypt encryption key.
Speaker 5: That data encryption key is, wrapped by And the key encryption key is what I'm using,
Speaker 4: you
Speaker 5: I'm using your password to create that. So what in this, mechanism, you you can, change your password without having to re encrypt all the data. It's it's happens, like, in a second. So that that's how you solve that problem.
Speaker 2: How do you know what what where to find the box? How advanced is this in the product?
Speaker 5: You mean, like, the location of the box?
Speaker 2: Yes. Yes.
Speaker 5: So right now, it's just based on labels.
Speaker 2: So it just
Speaker 5: tells you it kinda reminds you that, oh, this is where it is. And if you, like I was trying to show, if you have label on the box, you can see that, okay, it's in this box. 1 of the others what you can do is basically, like, if you have if you're packing a box, like, if you're moving and you're packing boxes, right, you can, take photos multiple photos of the same box as you're packing it. Right? So it Taylor a little bit of organization, but not quite that much.
Speaker 5: I know people who buy, QR code stickers from Amazon, put them on the box, link them to a, Wordle Doc where they actually enter things that are in that box. So if they need to find it, they can scan it and, you know, find things. But this is makes it a little bit simpler, but it doesn't have any location tracking or
Speaker 2: Is that your toy project or you're planning Two, to bring into Two to advance
Speaker 5: It started as a toy project, and then I in I, actually, I worked with my son on this. He's at UNC now. So we worked on this together, actually. So it it was kind of an, experience for both of us. Me trying to productionize something and, you know, him learning, you know, the system design and all these things.
Speaker 2: Yes. We'll have to have people talk to you afterwards.
Speaker 5: Yes. Thank you very much, guys.
Speaker 0: Alright.
Speaker 5: I'll give you this first. Still
Speaker 2: perfect.
Speaker 7: Can you guy you guys can hear me? Is it can you hear me through this thing? Okay. I'm Taylor, Partners. My, project is a, I Solves the Wordle, and then 3 LMS compete to build me a new website and launch it automatically.
Speaker 7: I will show you. This is the current 1 is I LMs are bad at what they're what they do. So I have my phone here because the dumb part of this is I'm gonna try to solve the Wordle in front of Avi right now. So, if that doesn't work, then we will just pick a pick another thing. I'll myself 1 minute.
Speaker 7: And screen mirroring isn't working for here, but, let me just see if I can do it. It'd be way more stressful if I had to do it up there. I got I got I got the e. Point no. Alright.
Speaker 7: I 3 I'm 3 way, but I wanna just kinda show. So, once I solve the Wordle, if I was better at solving the wordle, I have a shortcut on I, phone. Let me actually, like, bring up the architecture for this. So, it's okay. I, I have a shortcut on my phone, that has a, saved SSH key, in in my phone, and it connects through a tail scale to my home computer where all this is running.
Speaker 7: And I'll tell you why that in a second. So I that like a good we'll do horse. Alright. That's probably gonna make something weird. So right now, it is connecting with my computer at home.
Speaker 7: I have this, I don't know if anybody's using I, n t f y dot s, s h. It's like a push notification platform. And so I just got a push that said it started, which means that my computer is at home, running those commands right now. So there's a prompt. There's just a big Lloyd prompt that says you're a web designer in a competition against 2 other AI teams.
Speaker 7: These are the actual, 3 different LLM prompts, that get sent. So there is an orchestrator that's at home right now, that starts these 3 parallel prompts. And these all use my, the subscription plan. So there's no SDK stuff here. Like, I'm sure I could use the SDK, but it's using Lloyd, subscription, OpenAI subscription, and a Gemini subscription.
Speaker 7: And I know I don't need all of those, but work pays for them. So, we have this orchestrator that's going to tell all the different LLMs to make a I, but then it's not gonna know when they're done because it's just an LLM, and there's really this New super in-depth. So this, orchestrator LLM is watching a file folder, where these LLMs output their single page app Two. Not really an app, just their single page website. At some point, in line prompt, we got the spawning 3 competitors, via Claw, Gemini, and Codex.
Speaker 7: So now, we're Way. We we're not gonna wait for this. Like, we this might finish. Sometimes it's 5 minutes. Sometimes it's 11 minutes.
Speaker 7: Codex is the slowest 1 every single time. It's also yeah. It also does a pretty good job. Once that happens, I've given it a prompt of, like, my tastes. So there's another LLM judge at the end when everything is done that picks the winner.
Speaker 7: And then and then what will happen is once the judge finishes, there's a validator, because a lot of times these things maybe don't put their exact HTML in, and I don't want it pushing HTML that's not gonna work. So there's a little validator script that fixes. Again, another sub process. Right? Like, lots of, like, little mini agents doing little work for me.
Speaker 7: It updates the, app manifest like, the website manifest, which is the way that the, this is just really an HTML page running on CloudFlare where the manifest can see where all of the archived ones have been. It builds me a candidate's page. I'll show you all these candidates page in a second. There's an option where if if, I kinda look at them 1 I look at them every day because I enjoy doing this. I can pick a different 1, and I'll show you how that kind of works because it is an HTML page.
Speaker 7: So there's no, like, real database behind this. It's just an HTML page and some little Lloyd cloud CloudFlare magic, so I can override the winner. There's a predeploy safety gate that just makes sure that, like, it pushes something for today. It's not been failing very, very often. So I had it run 10 times yesterday by itself so for this demo, so I knew it would work.
Speaker 7: So we'll see if it's actually working. It's still on it's still on creation at home. And then it commits it using my, again, local GitHub, pushes it to GitHub. Cloudflare sees that, picks up the deploy site and deploys it. And then it sends me another little push notification, that gets to it.
Speaker 7: So, like, let me show you, the site. This was just testing. I don't remember what this 1 was exactly. This 1, I was this may be sunny. I don't know.
Speaker 7: This 1 makes no sense. So, like, if we look at some of the archive ones so I built this archive page so you can see that I a lot of times, the LMs don't make great I, but sometimes they do. And you can see that those little red circles at the bottom, these are my favorites. And the way I'm doing favorites is there's a Wordle, I have a token that I can put in the URL. This all sounds completely insane, and I know that.
Speaker 7: It's cool, though. Like, it's all it's all free for me. I have a token, on my on this computer that knows that it's me, so you can't just come in and favorite different ones. That sends a little, payload to CloudFlare, and I'm using CloudFlare, a kvworker. And it just, like, puts a little thing in the Hey Taylor favorites website, thing.
Speaker 7: And then and then that's how it's loading these from CloudFlare for just me. So, like, recently, this 1, I loved it so much. I don't remember what it was. Each 1 has a creative direction before it spawns the 3 workers. And so you see, like, ads, like, little little stupid stupid animations, brings you to the archive.
Speaker 7: Couple other favorite ones. This 1 is pretty cool. But you can go to the candidates for each 100, and you can see that, again, this is a really these were bad ones today. Claude, Gemini, Codex, and, it kinda has, like, a whole history. So this is all, like, on my website.
Speaker 7: You can go mess with all this. I don't think you should, but this was an this was a fun 1. This is I walking upstairs for the different for my different projects. And
Speaker 3: So these are all AI generated?
Speaker 2: These are
Speaker 7: all yeah. I don't I didn't like, every website here is AI generated. So, I think it's worth, like, talking about the prompt for just 1 second, and that is on the system architecture. Recently, I, a cup like, a month ago, I was like, they're all boring, and they're just CSS and HTML. So I allowed it to have, external resources only.
Speaker 7: So if if you can grab some resource from a CDM I Google Founded, 3 j s, p 5, you can use it. I'm not loading that stuff for you, LLM. But, like, if you can grab it, you're allowed to use it. But, you know, don't make a React app because I'm not it's on CloudFlare free, and I don't wanna pay Vercel or something to, like, host a stupid website. So there's no external links besides the architecture page that I made yesterday for here, which has all of the different things on it.
Speaker 7: I had Two add some, like, pretty Loop, you know, there's there's a lot of weird words when an LLM makes stuff. So, any any word that I was like, maybe someone doesn't know what this exactly means. There is a, just there's I a a hover over that. It will help you understand it. And then today, since I, like, had 5 minutes and I was bored, I made an architecture of the architecture page.
Speaker 7: And so it just says how this page is built. And And it's like a little Pendo tour, about how this Way just I how the whole page is built. Silly, silly stuff. Like, sometimes, well, I just gonna be fun to play with, like, instead of, like, working with them all the I, building, like, real real stuff. And so the horse one's not done yet, and it might it'll it'll definitely finish in the next, like, 10 minutes, but, it'll push here.
Speaker 7: Yeah. So that's my that's my little website builder. Anybody have any any questions?
Speaker 5: You said the they are
Speaker 6: given a direction where at what step is that?
Speaker 7: Yeah. That, that's at that's at the first that's at the first step. So, the orchestrator, and exit tour. At, the generate script has a a prompt of AI. I don't think it's actually on here that, takes the seed word from the from my Wordle answer and generates a direction and then gives that to all 3 of the different models.
Speaker 2: So that's
Speaker 8: the connection.
Speaker 7: Yeah. That's the connection of the world. So, like, I don't know what it would build for horse, but, like, in 10 minutes, if you go to heytaylor.dev, like, you'll see what it built for horse, and then you can see the candidates. And yeah. Yeah.
Speaker 7: So cool. Any other questions? Yeah.
Speaker 8: What l m for the l m as a judge?
Speaker 7: I'm using Lloyd for the orchestrator and the judge. Yeah. It it it doesn't pick itself very often because, like, honestly, the Claude ones are not good. Like, the clogged ones are very often, I, this one's kind of interesting, but, like, this was this was tweet this was my 1 of my favorites. We've again, weird stuff.
Speaker 7: Like, I don't know I don't you know? Again, Claude Gemini codex. They change a lot. Yeah. The model the models change a lot, and then my website gets better or worse.
Speaker 7: So, cool. Yeah. Thanks, y'all.
Speaker 3: I. Sweet.
Speaker 6: How's this working? This audio okay? Cool. So I wanna show off a project I've been working on for a couple of years. It's been an experiment that I've been running about using, using, LMs or VLMs to do data extraction for PCB design.
Speaker 6: And I wanted to show off kind of where this came from. So I I've done both as a hobby and professionally done some hardware design. And all of the softwares out there, these are like big multi thousand dollar software packages, that build PCBs and design electronics, all of this runs off of a data set that, essentially, you have to build and maintain as a I of the hardware. Two you're out there collecting all your little DeepAgents. You're putting it on your board.
Speaker 6: You're bringing all the data over and you're trying to get into this piece of software. And then every time you wanna update it, you gotta go back to Two all these components and and kinda find out what's changed and bring them over. And the data format that the industry uses to actually share this is a PDF. So, like, basically, this is where all of the structured information from all the manufacturers are bringing this data over in, you know, in a fairly standardized format that they put it in, but in an incredibly labor intensive process to get it out of here. So I've been really curious for a couple years, like, when were helms gonna get good enough to be able to pull this out automatically?
Speaker 6: I I've been pushing up I have a couple of these PDFs I've been pushing up to, like, Gemini and Lloyd and and and open and, and and OpenAI for a while to see what they could do. And they they're pretty good at it. The problem is they're incredibly expensive to do this with. So running up, you know, all of your, all of your PDFs through 1 of those, you know, frontier models gets really expensive for for design. So what I actually started doing over the winter, I realized that the VLM kinda category that's emerged Founded, like, New 3.5, 3.6 with these smaller VLM models.
Speaker 6: We're getting really good, and I was really curious to see how far we I push them to Two. And there's another category of models that's emerged over the last year or so, which are these really special purpose, OCR models, which are this, like, half billion parameter models. You could run them locally if you wanted Two, that are very, very fast. They're very cheap. They do a really good job of pulling markdown out and getting some structure.
Speaker 6: So I built an application, that in a pipeline behind it to kind of start to look at this as a problem and see if you can start to build a a kind of structured generation, pipeline that gets this stuff out. And this is a this is actually an app that I built that's in using the Zed framework. That's actually I had a talk with us through Zed. And this is actually using their, their UI framework to do this, all this document management. The reason why is it's actually a PDF viewer application.
Speaker 6: And over on the side using the little Zed UI, it's actually got a list of the extractions that I'm running. And these are all extractions that run through different models. You can start to compare them. So this is using a model that, 1 of the more interesting New that's out there right now. It's this really tiny, GLM model that's a half billion parameter, model that runs just OCR, just gives you markdown.
Speaker 6: It's incredibly cheap and fast to run. It does a decent job. So I went through and ran this entire 150 page, you know, PDF through it. Turns through it
Speaker 4: in a couple
Speaker 6: you know, on a on a GPU in the cloud. It turns through it in, you know, about 30 seconds, pulls all it out. And this gives a first pass that you can then use to start to build a kind of a hierarchy of the document. So this actually goes and extracts the type of contents out of the document. So you can have a roadmap of what's in there, kind of where things are laid out.
Speaker 6: And then the second part of this, and this is where I'm gonna push the live button and we'll have something break. This passes off to another model. I this is actually a quen 3 5 I, 30 I parameter drive model that's gonna generate a structured version of this document that has the kind of grounding for all the components and a much much more robust New of what's actually in those pages. And we'll see. There it goes.
Speaker 6: Came back. That worked. So this is another pass. If you I don't know if you can see on the I. There's actually the markdown extraction that occurred before, and then there's this next pass that came back from this model that's running QUIN QUINN 3 5.
Speaker 6: And what came out of this is actually all of the, you know, grounding and bounding boxes and much, much better I of, like, you know, description of what's in those things. That then can you can pass off to do, you know, larger component extraction from this. So what I ended up doing was, building this I'll show you a quick demo of what this does. This took the markdown data right here. It's gonna kick it off to a I model and it should come back right there with a structured generation of all the stuff in the document.
Speaker 6: I was just based on the markdown. The more interesting thing is that you run up the limits of what the markdown file can actually tell you. So you actually wanna go Two this 100 page document and pull out sections of it that are
Speaker 4: go ahead.
Speaker 3: Just then, was that the whole document?
Speaker 6: That was a that was a subset of the document. But it could've if it Way, like, it was, like, 10 pages, I think, it went through and did that. It was very fast. Yeah. So that's what thinking turned off on a Quen 3 6 model.
Speaker 6: I I'm running this on a GPU in in Wordle. So all this you're seeing right now is live on instances I had to put in modal. So they're, you know, they're not huge GPUs but they're like, you know, sized appropriately for the model they're running.
Speaker 3: Are you uploading Image.
Speaker 6: And the reason why this is interesting. There's some of these models use the g use the PDF content as well as the image. And you can kind of do both. But the problem with a lot of these things from manufacturers is they don't always give you the actual text version of PDFs. There's always an image.
Speaker 6: And a lot of I, the rendering of those is very complex and actually a lot of subtlety in that. So I decided to standardize on the image. And 1 of the reasons I, I'll show you in a second, this is what I want I wanna do next is and this is kind of where this goes from the using the markdown and the kind of, like, big chunk of the Architect, is you actually can go in and start to say, I wanna grab a section of this image because it care it contains something that I care about.
Speaker 2: Or it
Speaker 6: could be a whole page or a subsection of the image. And I wanna send it to a specific model. And I'll Way, I wanna get all of the JSON parameters for that section. This will kick off just a chunk of that image, send it up to the up to the thing and come back with a with a JSON object that contains that data. The reason why this matters is that there's 1 of the things I really wanna do with this is use it to fill in complex data that's in charts.
Speaker 6: And some of those require Raleigh specialized models or much bigger kind of frontier models to be able to do the reasoning on the chart. And you could pull out the chart. You can identify the chart. You'd already figured out where all of them are in the page with the bounding boxes. You pull those off.
Speaker 6: You could send those off to, you know, Gemini 3 1 or Lloyd. And and and you start to build this composite view of the document. So the cool thing about this, and this is I of the where why I did all this is, like, if you go over and look at the benchmarks, the first pass cost, less than a Way is it? Less than a it's I 3 it's like you could do Well, let me put it this way. Do a 100 page document for about 5¢ in the for in the in the markdown version on my GPU which is not very cheap.
Speaker 6: And then on the second page, the pass is about 6 and a half times more. So you could do that I on Architect basis to get the parts out that you need. You don't need to run the whole 100 pages. You bump that up to something like a frontier model. This is gonna start to cost dollars per document.
Speaker 6: So you can start to pull out the sections you care about and I of create this composite view. And so what's going on here in this thing, you can see this in the in the side here. This pane is really this composite view of this document with all these things stitched together. And when I started that little prompt thing, what what's really going on here is Way the beginning of a kind of an agent framework and start to reason about what it's already collected from different models and what it needs to send off to other models to get more fine grain, information. So it can build up that input that then gets handed off to Way, build me a Joseph object that describes this component.
Speaker 6: And 1 of the cool things I discovered along the way here is that I started out doing this as a cost savings measure. But when I started to compare the models across the board, what I what I found out was that none of the models were a 100%. There was even the the larger model would would have areas where the smaller model actually outperformed it on on certain pages for whatever reason. And when you start to run multiple passes with different models and you feed them all into 1, you know, kind of like summarizer, it can actually look at those and find when there's divergence and flag that and say, Loop, my confidence in this area is lower because 3 models didn't agree on this or 2 models didn't agree on this 1 here. These are all completely in agreement.
Speaker 6: Large and small model agree. That's a much higher confidence. So you get a you get a score as well as a, you know, a I of a second evaluation of that that same dataset. And that actually ended up working out really well from I of finding, you know, finding areas that need to be potentially even human intervention. Because what this could do is flag ones where none of the models could pull the table out or none of the models could pull the value out and a person could go back
Speaker 4: and check it. So go for it.
Speaker 8: Question about pass 1 and pass 2. Pass 1, j 1,
Speaker 6: pass 2. I so so I'll show you. So Way I did with with this is I actually started out using this is something I haven't been pulled into the in the table there. I started out using this Infiniti Parser Pro, which this is on the benchmarks that came out, this came
Speaker 4: out a couple of week
Speaker 6: or 2 ago. They claim that they are the best in terms of like VLM document extraction models. This is actually quen 3 5 with a bunch of fine tuning on top of it. What quen 3 6 came out right Founded the same time as this. It turns out quen 3 3 6 is actually better than this specialized model now because it's general purpose, but it's outperforming it.
Speaker 6: So I I did the the the initial Loop of this app and some of the data I collected was using this model, but I've actually switched over just using quen 3 6 and actually outperforms it. And these are both 30,000,000,000 parameter, like a 3 b models that are under the hood. Can I ask?
Speaker 4: Yeah.
Speaker 8: With the second pass, you get the grounding and everything.
Speaker 4: Yep. So then,
Speaker 8: how do you use the first, like
Speaker 6: So the first pass so the first pass Yeah. So you Way actually can if you can afford to generate it. So here's the thing. So if the first pass builds you table of contents, it gives you markdown of everything with include structured tables. And then Way you're gonna run into in that first pass is sections where you actually need to I pull out a table that's more complex than the the first pass could provide.
Speaker 6: You can grab that that table out of that page and do a post process just on the table. And then you can use that as like to augment it. Two it's almost like a zoom in on a detail. For most pages, you don't wanna do that though. You don't wanna spend the money on.
Speaker 6: I, a 150 pages, there's only, like, 10 in here that are relevant. And so that first pass gives you the full index. The table of contents tells you where, you know, where the structured areas are that are eventually relevant. And then the agent could use those to drill into the page that actually needs to go To refine. To I.
Speaker 6: Exactly.
Speaker 1: Can we get 1 more
Speaker 7: Yeah. Raviq, how how when you I are you testing the idea that you're doing a pass and you Two change the model? And, you know, are you manually doing that? Are you writing specific evals or just like
Speaker 6: I evals are a big part of this. So this is what I I created a corpus of documents. There's about 15 right now that are, like, kind of different types of components and different manufacturers. They're all using different formats. And I went through and I did a couple of the passes with those with a, frontier model Two to kinda create a a a starting point with an output to those.
Speaker 6: So you basically create an eval output from those. And then I actually went through and hand coded a bunch of the data Two I could go through and see where areas where none of them were doing quite right. These were ones that were hard pages. There's a number here that's hard to get out. And so I used those as a way to kind of run through that.
Speaker 6: And what I I up finding was that even with, like, the frontier models, there were some pages that were hard, but it was easy to find the divergence kind of measurement as a way to kinda drill in on those. But that doesn't have, like, Raleigh, really, precise, like, kind of quantitative view of those. It's more just I a heuristic for, like, this one's doing fairly well over here. This one's doing poorly over here.
Speaker 0: Good? Okay. Cool. I.
Speaker 6: Yep. I think I'm out of time. So sorry. Okay. Thanks so much.
Speaker 6: Alright. Thank you.
Speaker 5: Alright.
Speaker 2: I, Tim, I actually Lloyd him on Loop I you, if you wanna learn about, AI or or even persistent DeepAgents, he posts a lot of great content on, if you've heard about OpenClaw, a similar thing called Strix, I I think that's what he's talking about. Yeah. But definitely follow him on, Blue Sky if you if you have 1.
Speaker 3: I. Okay.
Speaker 2: Alright.
Speaker 1: Can you hear
Speaker 3: me? Yes. Cool. Yeah. So I'm talking about, OpenStrix.
Speaker 3: It's, like a stateful agent harness. This here is, is 1 of my agents that are just built as runs on OpenStrix. His name is Motley. If you have your laptop open, you can just, like, install this thing and get started if you want. I mean, it's a tinker.
Speaker 3: It's a Meetup, so might as well tinker. Right? But, Way, so, I have several of these agents. Basically, if you think of if you know if you're familiar with the open claw, it's kinda like that vibe. There's also, Hermes.
Speaker 3: Hermes DeepAgents also another 1. I kinda distinguish this class of agent from, like, Lloyd code or or codex because, I went down like this. See. For 1, these things tend to act ambiently. So, like, instead of just simply reacting to, like, a message you send them, it's also reacting to, like, you know, something that happens out in the Internet somewhere.
Speaker 3: You know? They're always kind of awake and aware and and looking at things. They're also, let's see. I, I would say oh, well, state. The state.
Speaker 3: I guess the state's the important thing. I'd also argue that they're I of more more general. So, like, codex and Lloyd Code, they advertise to Two, like, coding agents, but they're actually, like, more general than that. You can use them for a lot of different things. These harnesses, they're advertised to be to work generally.
Speaker 3: You can you can like, I I have 1, I use at work for coding, and I have another person at work who's using it for, like, Architect, and it works fine. Right? But as you begin to use it, the harness might be general, but, after, over time, it comes very narrow real really quickly. Like, this guy, Motley, he doesn't do any work. He's a he's a joker, and he just makes fun of people.
Speaker 3: So, I mean and then, here's here's, the the the namesake. This is actually came long before OpenStrix. He's does, like, AI research or something. But here's Motley's Way page. If you wanted to go here, he has some really fun games he's made, like the, the pre trained remain.
Speaker 3: You can, like, pretend like you're an LMS, see if you can go through the pre training of guessing the next Wordle. And, anyway, it's just whatever. That's his own thing. He's a little goofball. So Way, I have a whole bunch of these things.
Speaker 3: I was gonna walk through the code a little bit, but not too much. I I don't know. I I can go a lot of different places, but I don't have much time. So I think 1 1 way this this is kind of a unique harness is I kinda incorporate a lot of cybernetics principles, from, like, the, vital systems model. I'll walk through that real quickly.
Speaker 3: Yeah. I don't know if this is
Speaker 4: a really good way to
Speaker 3: do it, but, so I have STRIX generate this image. But, basically, the the viable systems model, like, it's it's like this this New. Like, first of all, it's hierarchical. So, like, you are all individually, vital systems, and then you're, like, us all in a room. We're all 1 vital system.
Speaker 3: Each, like, organ in your body is a vital system, and each cell is a vital it's just, like, this Wordle hierarchical nature to it. The guy who came up with this back in the seventies, he did actually to to understand, like, you know, organizations. And so, it it works really well for understanding agents as well because, like, you you like, your agent is a vital system, you together is a vital system, and then, like, how your agent fits and, functions on on your team, like, if you're doing software engineering, for instance. It's also, you know, I viable system. But then he also breaks it down into, like, different, like, systems.
Speaker 3: And so the way I I think about it, like, there's 5 systems. I don't wanna go too deep into it. Operations, that's boring. It's just I what you the agent does. Identity and policies is interesting, but it's it's I spend most of my time thinking about, like, s 2 s 4.
Speaker 3: So, like, s 2 is, like, coordination and, like, how the agent, like, resolves, like, conflicts with, other other, you know, individual systems. So I think of that as, like, more I looking at siblings in the hierarchy or or inside. Then f score is, like, looking upwards. So it's I it says Two. If you think about it I military intelligence, business intelligence, where it's, like, looking out in the world and seeing what's out there.
Speaker 3: So it's it's, it's kinda how the the I system or the agent gets sewn into, like, your social structure, I guess, in
Speaker 9: a way.
Speaker 3: And then s 3 is, like, resource allocation. So, like, if you're trying to plan work, it's that's I of the whole point of it. It's figuring out, yeah, x amount of time, how will I do? So that's so I don't really code that directly into the agent. It's actually kind of embedded into skills.
Speaker 3: And in in the onboarding process, like, if you were actually to go run that command and actually install, OpenStrix, it kinda walks you through, like, what
Speaker 4: do you wanna use this for, what
Speaker 3: do you wanna Two, and it'll set up some processes. I, oh, you wanna, you know, watch your GitHub repository for issues I in, which it sticks manages most of the issues that come in. It'll go up and set up, like, a Taylor to make sure, like, it's listening for those. It doesn't involve the ALM every time. It waits until there's actually something to get changed, and then it involves the ALM.
Speaker 6: Or we said Wordle set
Speaker 3: up all these things. So, like, something like that, I, a issue coming in, I think of that as, like, being escort, like, looking out at the world and kinda and and seeing what's what's out out there versus, things with, with Kiel, my work agent that I use, which she's basically, like, a fully autonomous software engineer at this point. I I designed that this morning. Because, he'll he'll,
Speaker 0: like, he
Speaker 3: has, like, pullers for checking on, like, if there's a certain file that got modified in the the repository. There's, like, I gotta make sure it's in sync with this other I, so we'll just automatically pull it down and, like, make the changes and submit the pull request, without, like, I really being involved. But so there's a there's an onboarding skill that is kind of invoked when you first start after that, and the agent shuts it off when the plug is done with that phase. And then I have a whole bunch of other New, like the introspection skill. There's just I of a whole bunch of, like, files inside for just instructing, like, how to how to understand if there's a mistake made.
Speaker 3: Or if if you, the user, asks, why would you do that thing? It'll go off, and it'll Wordle tell it'll tell, the agent, like, how to look at its log files to interpret and actually answer realistically what it actually did. I, yeah, so there's all these skills. I would say the code itself is actually not that big. There's, like, maybe 10 files, which is feels
Speaker 4: kind of Lloyd, maybe.
Speaker 3: I feel like there's probably more English than there is I, but that's kinda how it goes these days.
Speaker 2: Skills, are they transferrable between different agent harnesses, or are they I strict specific?
Speaker 4: Most
Speaker 3: you can plug most skills into this, and it would work. These are probably all specific to the Partners, because
Speaker 2: I think
Speaker 3: that's true. New. Actually, so prediction review, this 1 is probably not. That one's interesting. Like, I Two this idea from, like, psychology.
Speaker 3: Like, Carl Young has this concept of, like, you can't like, when when you're talking to a patient and it's, like, a therapy situation, they usually lie to you because that's, you know, what you do in your therapy, I guess. So to get around to that, the the the psychologist do they'll, they'll, like, make a prediction I a patient. So I think this patient's state is in this state. So if that's true, then I make this 1 change, then I should expect this result. Right?
Speaker 3: So, like, did they get away from they get away with, like hey. We have, like, LLMs hallucinating all the time, so they, you know, I to us. Two the same sort of situation, the way you get around it is you just make predictions about the future, then you come back and, see if if it actually you know, you get the right
Speaker 0: method model.
Speaker 3: And so that particular, skill, I feel like is has been very useful. Like, I've had, like, a lot of a lot of the agents start to, like, learn. I get a Centennial model about how I work, how I work on my team, and how, like, different team members where they where they function.
Speaker 2: You can
Speaker 3: start, like, making predictions, and if they're wrong, it'll actually correct their their memory. So this
Speaker 6: is pretty
Speaker 3: cool stuff. 2 minutes is not enough time to go deeper things, so, I guess I'll just do questions. I don't know. Anything else?
Speaker 2: So for the prediction so it's predicting both yours your state and also the agent's own state?
Speaker 3: Anything, really. It it's, so the Motley, the funny 1, I have him, predicting whether or not his jokes were funny and, like, people, like, laugh at him or not. I I don't know what Kiel does. And Strix, let's see. I I think he predicts, I don't know.
Speaker 3: It's smarter things. I'm not sure.
Speaker 1: Are you are you running all
Speaker 2: these in, like, a Docker container or a cloud environment?
Speaker 3: Or Good question. So the personal DeepAgents, so Motley and Strix, they they run on a bare VM next to each other. I feel like I should put them in containers. They're they're kinda getting to that point. Then the the work agent I have running on my laptop, and I just keep it so it doesn't go to sleep and I'll work all night Consulting things.
Speaker 3: So
Speaker 2: So that 1, theoretically, can access your files and all that stuff? Yeah.
Speaker 4: I you guys hear me? Okay. Oh. 0. It's got an adapter.
Speaker 4: You need an SBC?
Speaker 8: Maybe. Yeah. I think so. Yeah. So
Speaker 6: we got you here.
Speaker 3: Okay.
Speaker 8: Hey, guys. My name is Zachary Lloyd. My background is maybe a little bit different than a lot of people here and that I'm actually a corporate Taylor, but I like Two, code and build with AI in my spare time. And I've started playing around with, some of these tools, to try to build something that I can put to use in my, in my actual legal practice. So I wanted to walk through, kind of my first foray.
Speaker 8: I call it my first foray into trying to turn Claude into a contract lawyer. So 1 of the things that I'm running into, 1 of the issues that that I I'm running into is with context, because a lot of the, contracts that that I'm, working on are are extremely long. I have a lot of them. I I like the ability to just kinda quickly, you know, access the contents of them, and I be able to do it accurately. And I find that sometimes, you know, if you're just attaching contracts to your prompt, you know, or or just kind of trying to Two, like, a a basic search, a LLM search with a, an LLM over a database of contracts, it kind of starts to break down at scale.
Speaker 8: So I started playing around with a potential solution to that. And so Way I I'll go ahead and submit this, question. Way I started this is what I started with. This is not the final product, but
Speaker 2: I wanted to show you
Speaker 8: what I started with I I thought it would be interesting to see the show the evolution. So, this is just I a simple little RAG agent that I built to search over a, sample database of contracts. It's called the CUAD database and, it's publicly available. It's about 510 contracts. And, I tested this over 5 different sequences of questions.
Speaker 8: Got some decent results, but you're probably gonna see it takes so I'm running QIN 3 32 b locally, under the hood, so it's gonna take it a minute to respond. But as you can see, it says, I don't have information about that. So this is kind of what I'm talking about that like I you know, these LLMs sometimes struggle a little bit with the, you know, the I, when things get bigger, they, I'm not sure exactly what's happening there. Could be something with, the chunking strategy. I'm using a pretty naive chunking strategy right now.
Speaker 8: Could be something with the Querying layer. And so I started thinking about how to fix that and looking into those things. But as I was doing that, I went to the All Things AI Conference and I heard people talking about this thing called a Model Context Protocol, which for me was very new. And so I started researching that and I thought, you know, Way, maybe before I start fiddling around with all the code under the hood, maybe try wiring this up as an MCP server and see what happens. And so that's what I did and I'll show you this and hopefully this, this usually gives a pretty good answer here.
Speaker 8: So I'll ask the same question. It's wired up as an MCP server. I've added in the tool layer, and, yeah. So it's loading the tools. It'll take it a second to answer.
Speaker 8: But typically, with the MTP server, no other changes, same chunking, same querying strategy. I'm using Chroma DB for the embedding. And, it it just works a lot better. I I found that really interesting that, I could get, like, a I mean, the the 5 question sequences that I I mentioned testing it with, pretty much the MCP server handles pretty much all of them. I don't get any really wonky answers.
Speaker 8: You can see here, it's giving me the this is the answer it always gives and I verified this is the correct answer for that contract. So, just by Consulting my RAG DeepAgents and wiring it up as an MCP server, it started working better for me. So then I started deciding, like, I'm I'm using this right now with a sample contract database, so it's not, like, really that useful though still. So I wanted to think about, like, how can I actually turn this into something that I might, you know, if I wanted to decided to actually wire up, and use in my practice? And so, the next thing I did was add a layer of code for ingesting other contracts so that, like, a a user could, point the ingestion code at their, at a folder of their contracts and create their own collection and then have, have Claude I this MCP answer questions about their own database of contracts.
Speaker 8: So I did that. Right now, that just, runs via the, CLI. It's, yeah, just a New CLI DeepAgents, but I'll just show you an example. So this is, this is a collection I had already created, of, publicly available contracts from Silicon Valley Bank. But you can see so you just run that and it, chunks them, creates the Millennium.
Speaker 8: And then I will have to restart Claude because I rebuilt that collection. But once you do that, you can, ask it questions about, your own database of contracts. But I'll also show you, before I get to that, I'll take you through a couple of previous chats. So when I first did this, it didn't work so well. So you can see, I tried to ask it about, Silicon Valley Bank contract.
Speaker 8: And again, it says, I wasn't able to find it. It's still, if you look at the beginning of the chat, it's still answering that
Speaker 2: initial question correctly, but
Speaker 8: it can't find Two, I, but it can't find the, contract in the new collection, the custom collection that I created. So I started looking into why that was. And, I found that it it sounds like with MCPs, your doc strings and your tool signatures are very important. So, the first thing I tried was adding a sentence to the doc strings of the Two, basically saying, like, if you can't find a contract initially, search the other collections. And so that yielded a marginal improvement that you can see here where it is still saying that it can't find the agreement.
Speaker 8: But at the at at least at the end, it's saying, like, do you want me to list the other collections? So it's still not really working, but, it's at least, like, kind of saying, like, I know that there are I can I understand that there are other collections that you've created of contracts that and, I'll I'll list them if you want me to, but that's not ideal? So the next thing that I did and what it eventually fixed it is that I made it optional, initially Two, give your collection a custom name. If you don't specify that in the command line, it just gives it, like it just puts all your contracts in, like, 100, collection called legal contracts. And in line with that, with the MCP, I made collection name an optional parameter, and that was creating issues.
Speaker 8: And so when I changed it to a required parameter, we'll go back into the original conversation and try this question. And, again, it's gonna take you to second to answer, but, usually, it's able to get the correct answer to this question. So that that's how I that's kind
Speaker 2: of where
Speaker 8: I'm at now.
Speaker 6: Can I ask
Speaker 2: a question?
Speaker 7: Yeah.
Speaker 2: Is the empty piece of a wrapping a keyword search or a vector search, or what is it what is the what is it providing to the to the AI?
Speaker 8: So it's basically providing, I think it would be called a vector search. So it's, it's creating, like, basically a a collection with Chroma DB. And then, the, underlying code, like, queries that collection. It, like, takes the prompt and queries that collection and returns the 10 most relevant chunks of contracts, and that's what the, that's what the host l l LLM gets access to. And so you can see it answers, that question correctly, and it even provides a little comparison to the, the original SunTron agreement that I asked about.
Speaker 8: So, that that's I where I'm at now. The the next, I think, kinda step that I wanna take, and maybe I've I think I've got a little more I, so I'll I'll just run this question as well. So this this is the type of question as a transactional lawyer that I really want to get, my LLM comfortable with answering. So this is more of a general it's not about a specific contract. It's asking, if there has Silicon Valley Bank ever agreed to a dollar threshold for cross defaults?
Speaker 8: And if so, what amount is typical? Really common question that you see in transactional law. You'll have opposing counsel make a comment on a contract, and you'll be like, well, I remember like, I think we saw that situation, like, 7 months ago, but I can't remember what contract it was from or how we handled it. And so being able to have a query like this answer correctly would be hugely valuable to me. I've been getting spottier results with this.
Speaker 8: You can see it's taking some time because it's, like, running through the entire, database of contracts even though it's pretty small, this, collection of SBV contracts. So that's kind of a a next step that I wanna look into.
Speaker 4: I have a question.
Speaker 6: So if we as a lawyer, when you're looking at this document, would you know exactly where to find the termination contract in the document without having reviewed it already, or is there, like, a obvious place or structure that's gonna be used in each of these that'll be consistent, or is it gonna really vary by the plan who wrote it, the format they chose, the the structure of that?
Speaker 8: It's generally gonna vary, you know, across I. Like, for the same client, the same client might if they have especially if they have some negotiating leverage, they might be able to use their paper for all the contracts, and then, of course, it's gonna be in the same spot. But across clients and and even within the same client, it's it's often gonna be in different spots. So, yeah, it's it's still looking. It takes it a while to answer this 1.
Speaker 8: So I yeah. I'll go ahead and cut it, but, because I know I'm overtime. But, yeah, I appreciate you guys having me, and I'd I'd love to hear any feedback you have.
Speaker 2: And you're here for (PostHog’s. Right? Yes.
Speaker 7: You weren't sure?
Speaker 2: Okay. So Ryan's with our sponsor, post hog. I think I mentioned they have a really cool website, so definitely check that out.
Speaker 9: Yeah. And I brought some hats and stickers if anybody's familiar with post hog or wants to post hog stuff.
Speaker 2: So that's why on the vote page, he's because he has a sponsor, he's not eligible to Way, but, I think
Speaker 7: let's see.
Speaker 9: This 1 is much more basic than a lot of these, and that's I
Speaker 4: the point. Let's see. Do that work. Here we go.
Speaker 3: Okay.
Speaker 9: I all of you, I've been building a lot of things, mostly I developer tools and stuff that nobody uses. And so, I recently built, an I clone just literally to use. This is the home page for it. It's it's I couldn't be more basic of an app. You can see just the comparison here.
Speaker 9: That's an Evite, invitation on the left. There's a bunch of stuff, and it's really hard to tell what you're invited to. And then this is what my invitation Loop like. I literally made this for myself. I did a little SEO, and a couple 100 people found it really quickly and started using it.
Speaker 9: And that was really exciting. And so then I was like, well, I work at (PostHog’s, I should probably start tracking it. And so I started, a little bit late, but these are I of some of the metrics around it. The people signing up, creating events. And so I found a couple of things.
Speaker 9: 1, you know, we all are, like, excited about building things, but we all suck at distributing things. And so, like, this was kind of that neat thing where it's, like, super simple. People aren't gonna go, you know, I code their own Evite. They're gonna just use something they can find. But also, like, so it self distributes.
Speaker 9: Right? Someone gets invited. A bunch of people now see I, oh, what's this thing? And they go and sign up and try it. And so it works really well for distributing it.
Speaker 9: But what I found is that a lot of people were I giving up on it really quickly. And so I've never been a fan of session replays. I don't know if anyone's familiar with session replays. They're kinda helpful. They're kinda cool.
Speaker 9: Basically, just shows you how people are using your stuff. This is not a (PostHog’s shill, by the way, but it is using post hoc. But basically, these are really helpful when you're I, why are people not using the app? You can actually these are real recordings, so sorry for showing people stuff. But these are people I getting invitations, responding to events.
Speaker 9: And so I this is what I was looking at and then like feeding the results into Lloyd Code or whatever else I was using. And so that was I helpful for a little Wordle, but then when you get thousands of people using it, which I accidentally did, I couldn't watch all the session replays. Session replays are actually not recordings. They're DOM snapshots. And then, we have a fork of our web that we use.
Speaker 9: Basically, it's all just JSON changes. And machines are really good at reading JSON. And so then I was just dumping all the JSON into cloud and I, Way, bunch of people are doing this. Can you help me fix it? And so we actually ended up building this into (PostHog’s.
Speaker 9: So we have our own DeepAgents orchestrator called post hoc code. And so, these are all the agents that are running currently. They're not connecting right now because the I here. But, this actually we have this inbox. And, it's actually proactively watching all of these session replays.
Speaker 9: Finding like, hey, you got a bunch of people that can't use your date picker on mobile because you I coded this on your desktop like a dummy. And so, like, it'll go find those things, surface them for me based on just, you know, tons and tons of JSON, and then show that to me of I, hey, we need to fix this. You know, I can go in here and just literally, like, kick off a task, and Way, like, yeah. I wanna fix this. It'll it'll jump into my, all the agents that are running.
Speaker 9: And then this actually runs around the clock. So, like, I do go in and, like, select some of these and work them. Sometimes I just wake up and there's PRs ready of, like, hey. This has changed. Second issue there is, like, Way, the campus– is just shipping a bunch of code for me.
Speaker 9: How do I know this is correct? And so the next step that I took with it, again, not (PostHog’s shilling, but a little post hoc Meetup, is, I added a tag in my CI workflow. So if I add this UI review tag, it actually kicks off, spins up the app locally, does a couple I scripts through it, and then it comes back and comments on the PR those flows. And so I can actually back in (PostHog’s, see my CI bot, checking on this RSVP that it created locally in my CI flow and make sure that the UI is working like it wants to. So, yeah, that's it.
Speaker 9: This is what I've been working on. It's been really exciting to have people using something. It's, like, kind of crazy watching people use a thing. So this is Way, I've been working on. That's it.
Speaker 9: Very basic.
Speaker 6: Like, kind of failure modes or is it like, are you like some sort
Speaker 3: of like meta structure using kind of group things together or is it like case by case you kind of find out what's going
Speaker 9: on? Like finding out how folks are struggling with it? I. Yeah. So I mean, a lot of that is through I mean, traditional like product analytics of I looking at but again, like, you see a drop off.
Speaker 9: It's like, I don't New. Like, everyone stopped using it. What is it? So the session replays have been like the magic unlock for that, and I don't watch them anymore. Like the only ones I watch are the CI ones.
Speaker 0: When it when it goes through, does it have a
Speaker 6: kind of way to pin them into different categories? Yeah. Yeah. Yeah.
Speaker 9: So in the agent orchestrator, it'll actually find, those common patterns. So I in this particular signal, these are all references to individual replays that have similar patterns in it. So it obviously, like, from a high level, it does aggregate patterns. It's not just like, hey, someone, like, had a problem. It's probably a browser plugin issue.
Speaker 9: But it does a really good job of it basically summarizes each of those sessions individually and then, finds nearest neighbors to those and then, like,
Speaker 4: surfaces those. I mean, I don't know
Speaker 2: if this is, like, 1 of
Speaker 4: the coolest, like, debugging things I've found
Speaker 9: across the board.
Speaker 6: I've found across the board.
Speaker 9: Yeah. It's been really cool.
Speaker 6: Log file. I mean, anything you, like, categorize the failure modes and, like Yeah. It's really amazing.
Speaker 9: Yeah. And it's all I mean, I have it's all it's all connected together. So the the logs are in here, the air tracking, and the replays, and, like, all the events. And so it really does a good job at, like, finding stuff that I I never would. And then the other challenge is like the majority of the user base are people that other folks have invited.
Speaker 9: Two like my users will tell me when something's broken, but the I 100 people they invite to their event, like I can't RSVP and they just like abandoned. And so like those people are never gonna talk to me. So it's it's really helpful when you have like a user base that's not actually your users, which has been really cool. So
Speaker 0: I I I remember I was at a search engine optimization event.
Speaker 1: And I remember you
Speaker 0: were talking I, like, when people are dropping off on e I sites. Yeah. And it makes me think about that for use case of, like, the people maybe added to a cart but never finished the purchase.
Speaker 2: Yeah.
Speaker 9: It's a similar idea of, like
Speaker 0: off and it's
Speaker 9: it's like you can see the signal, but it's like, but why? And so like understanding like, oh, they can't push this button or like this form is broken. It's not like throwing an air. They're not gonna, you know, call you and say I I can't check out, but yeah. So kind of now so now my mind I just like, I, what are other I things
Speaker 2: I can just, like, rip off and try to see
Speaker 9: if people use it?
Speaker 2: Example of something you didn't expect that was really surprising from the friction of the UI?
Speaker 9: The biggest thing is, like, people never use software how you think they're gonna use it. So like you build a cool date picker and it's like this is not how people wanna pick dates. Like most events are single day events, so like why do you have a calendar for the start date and the end date? Like it's just literally that. And so like that was a huge abandoned flow that again, I never would have figured out is I, people are getting confused of like, why do I need an end like a different Way.
Speaker 9: I just need a different time. And so I, things like that. The other thing is like, I did get a bunch of false signals early because we have the concept that I built into like my product analytics of like rage clicks. So like someone's clicking a bunch on something and nothing's happening. But I didn't take into account that like, especially on a mobile device, when you're changing a date picker, it looks like I a rage click.
Speaker 9: So I'm like, man, everybody's like really mad at this form, but it's like because you're like toggling over weeks or or months or something like that. And so there was a I did have to Two weed out a bunch of I false signals like that and and kind of trim how I'm collecting data for that. So, yeah.
Speaker 7: Wordle I the token usage for going through, you know, hundreds and hundreds of I, of these I multiple, you know, like dumb screenshots?
Speaker 9: Yeah. Yeah. I mean like you my work pays for it. So, it's a little different, but I think I think last month, it's I spent probably $200 on this. So it's not crazy, because again, it's it's largely JSON.
Speaker 9: I, it just loads the DOM once, and then it's just JSON. Yeah. So, So
Speaker 7: does it work with the like, the mobile SDK replays?
Speaker 9: Not as well because those are campus– captures. But we're we're we're testing different models with that as well. So yeah.
Speaker 2: What are the little things walking?
Speaker 3: It's
Speaker 9: a hedgehog. You can, you can actually control him too. He can jump he can jump on certain Dom elements. It's it you can turn him off Two, but that's just the (PostHog’s thing. So cool.
Speaker 9: And I I said, I'm for post talk, so, not to shill it, but I do have hats and stickers or if anybody wants to chat afterwards, I'll be around. So thank you all.
Speaker 6: Alright.
Speaker 4: It's
Speaker 2: not working.
Speaker 4: Okay. There we go. Awesome. Linux screen screen sharing is always tricky.
Speaker 2: So We've got a handheld if you can just
Speaker 1: throw that.
Speaker 0: That's I.
Speaker 4: Is this working? Yeah. Yeah. Alright. Awesome.
Speaker 4: Yeah. So I I drove to the wrong place today as I was that's why I was late. And I drove to the original location. Downtown. Oh, it was in RTP.
Speaker 4: RTP. Yes. Like, an extra 30 minutes. Brutal. Missed all the pizza.
Speaker 4: But maybe I get to meet you some of you guys after. So alright. So last night, I asked my AI agent, we're SSH into a a computer at my house, and it's got Lloyd code running on it. And I asked it to write build a little app to share recipes. I?
Speaker 4: And so now, like, watch this. I'm gonna ask it to deploy the app to the Internet. I? Deploy the app to the Internet. And so it's gonna do some stuff, and it's gonna give me a URL that's publicly ratable that y'all can hit on your phone.
Speaker 4: And, yeah, I'll show you guys all about how it works, but let's see. Yeah. Go for it. Yo. Right?
Speaker 4: And so we we are conscious about security because, again, I'm publishing an app to the Internet that anybody can reach. And so, like, in in the skill, it, you know, tells the agent to, like, make sure the user knows what they're doing. So we got a URL here and a Wordle, like, QR code Two. But let me open the URL. We'll come back here if y'all wanna try I your phones.
Speaker 4: So this is the URL and it might take a second to resolve, But I'll start explaining how it works. Like, if y'all use video conferencing, I'm sure you Two, Teams or I do I say Teams? That's like the worst 1. Zoom or Meet or whatever. Right?
Speaker 4: Yeah. Yeah. So Webex doesn't even work there. And anyway, if anybody works for Microsoft or whatever, I apologize. I don't wanna hit Microsoft.
Speaker 4: Two actually establish it actually establishes a connection directly between your machine and your web conferencing peer, right? It's a p 2 p then. And because you wouldn't want to route, like, all that video traffic through. Anyway, so it occurred to me that we could do the same for web traffic. And there's a bunch of people now that are, like, running, you know, AI harnesses, you know, OpenClaus or whatever.
Speaker 4: Strix Way the other 1. They're running them on their on their, you know, partially open laptop overnight or, or I a mini Mac mini or something. Right? And, like, so you already invested in hardware and you invested invested in connectivity because you pay your ISP every month. Right?
Speaker 4: So, like, maybe you wanna host apps from there. Right? And, like, the things build apps for you. Right? So, so check this out.
Speaker 4: Right? So we got a little recipe app and we can make 1 and stuff, but I'm just gonna get into the text. So you can see it's it's pulling JavaScript and stuff like any like any app would, but notice. Hang on. So it says service worker here.
Speaker 4: Right? So it got routed through like, the way that this thing works is, the app comes up, it establishes WebRTC between, like, the the computer in my in my basement and the web browser. And then the the service Wordle, which runs here on the browser, intercepts every fetch request, that the application makes, and it routes it through WebRTC instead of trying to hit something directly. Right? And so, like, that gives you basically peer to peer communication.
Speaker 4: So if you're into, like, WebRTC, Chrome has a WebRTC Centennial thing. And so this is the 1 that we're just checking. And so this shows you, like, how it actually established like, this does it does not punch through I through. So, like, usually, your computer Consulting behind a router, and you need to you you can't just, like, hit that IP because you're outside I, you know, another network. So, like, here we are in this, you know, university network, and it's punching through the, the NAT here, and it's punching through the NAT in my house.
Speaker 4: And both the computers are talking directly to each other. So this is how it show this is what's showing, and there's, like, some cool graphs and stuff here for how it worked, requests and all that stuff. So we're basically piggybacking on it WebRTC for it. And that shows some code. So I mentioned I mentioned there's a service worker and I, and some JavaScript.
Speaker 4: So this code is kicking off the WebRTC handshake.
Speaker 2: And
Speaker 4: let's look at some other code. So then this this is the service worker itself that's intercepting the I requests. And then Where is
Speaker 3: it running at?
Speaker 4: The the service worker was on the browser.
Speaker 3: Where is the yeah.
Speaker 7: Where is the server running?
Speaker 3: The server's on the browser?
Speaker 4: No. No. No. No. So the server is in the computer in my house where the same place where, like, the AI agent is that wrote it.
Speaker 4: So, like, it would be, like, where the where your Strix guy is running. Right? This is Way the server runs. But it pipes up but it but it it pipes directly to the service worker on the on the browser. And so, like, my infrastructure like, I I do own, like, a real DNS.
Speaker 4: Like, if you look at so here, let's let's let's break down the URL. This might this might help. So this says recipes. That's the name of the app. This says odd moon 86 39.
Speaker 4: That's an assigned alias for your computer that you get. Like, when you run it, you get 1 of these. And then Two. It's like p2pier but p2pier.claw.com. Right?
Speaker 4: And so that this is a DNS that I that I own. And but so everything routes there just to make the handshake. Right? I'm kind of a broker. So like browser goes to Peter Klaus' I infrastructure and it says, and it gives it I the addresses that it should try to go WebRTC with, right?
Speaker 4: Then handshake happens and a service worker intercepts everything and routes over WebRTC. So, yeah. So like, I mean, another I, this it's kind of like if you guys saw, like, Andrei Carpathi, you just had a Two. Like, he did an app about like, he built a menu analysis thing. You take a picture, you know the menu you want.
Speaker 4: And he's complaining that, like, the hardest part wasn't writing the code. It was, like, actually deploying it on Vercel or whatever because you have to, like, go buy DNS and, like, click a bunch of stuff, and, like, it's all stuff you have to do by hand. So, like, this is an agent first way to deploy stuff, and you don't have to pay Way SaaS. Right?
Speaker 2: So is this like if you have a Vibe coded web
Speaker 4: app Right.
Speaker 2: Connected, you know, to host it publicly, do you just include this library or is it something
Speaker 4: It's not a library. Your app doesn't have to know about this at all. It's so skill that your, that your AI agent harness like needs. And it'll run a daemon on your host that does the I. Right?
Speaker 4: And then I'll send the browser this, like service worker WebRTC code. So like this is the service worker WebRTC code and, here's code running on my edge that is used to make to do the handshake. So here's the handler for for that. Like this is serving this is this is the bit that literally serves the the JavaScript that ends up on your browser. And and then I the the actual agent that runs on your computer, like the server, is called this box agent.
Speaker 4: So this guy, this is the thing that's like this stands up on your side and does the WebRTC handshake from over there. Right? And then I Founded basically funnels all the traffic. So it's TCP over WebRTC is basically what's going on. Yeah, so if you all have I, you know, a laptop that's gathering dust and you wanna serve apps, I feel like everybody here is tinkering and like trying to deploy stuff.
Speaker 4: You know, give it the This is the skill, right? So you could point your harness at you know, this a key to claw skill, and, you just start deploying stuff. Any questions? Yeah.
Speaker 3: Clause, x kinda I a proxy to, like, set up the connection. And then is it, like, in the middle of the whole time, or is it hand off Yeah.
Speaker 4: It's not. It's not. So, like, if you
Speaker 6: do so you can do, like,
Speaker 4: a tunnel like n g I. Right? And so that thing is, like, literally in the middle of all your traffic. That's more of, like, a proxy, if you will. I'm more of, like, a let me stand up and let me just like, let let me help you do your handshake for for WebRTC and then get out of the Way.
Speaker 4: And then you keep serving directly. Right? Like like like in Google Meet, like, you know, all the video traffic isn't going through Two Google's edge. It's going direct. So same deal.
Speaker 4: Now there are some weird cases, where I like for example, there's some weird cases. Like for example, when you have to I upload like in this app. In this app you have to upload a recipe. I? And so that piece is, like, something that's, like, hard to do via intercepting the service Wordle, and so I tunnel.
Speaker 4: And so that will something like that'll tunnel, like the edge cases. But for the most part, I'm out of I'm out of the middle. Yeah.
Speaker 6: And that's just because the WebRTC interface doesn't let you
Speaker 4: It's because the way the browser works.
Speaker 6: The way the browser has,
Speaker 4: like Yeah. The browser, like, the browser, like, doesn't actually do a fetch. It, like, directly uploads. Yeah. Then there's, like, a few gotchas.
Speaker 4: Like, there's a couple of, like, Chrome integrities I've run through I've run into in the last couple of weeks, like, testing stuff. I, another 1 is, like, is, like, you wanted to support basic auth. Right? It's like so that's a that's a good 1. Like, basic auth will, like, pop will will bring up a a pop up outside of in the in the browser I, and I can't get at it from the from the service worker.
Speaker 4: So so because of that so because of that, I intercept basic auth and, like, put up my own form. So if you have so if you do have a app that has basic auth, it will work. It will look a little different, but it'll still work. Another yeah. And then, like, another tricky thing would be, like, say you want your app to do, like, webhooks.
Speaker 4: Right? Like, you wanna you wanna your server to be able to to to take webhooks from, like, you know, some external service. Like, they're not running a web browser. Two they're not running a service worker. So, like, this path doesn't work.
Speaker 4: So for that, I have to tunnel. And I do. Yeah. So, I, I think I think it kinda covers most of the cases in which, like, you've got a server and you wanna put it out Two the Internet to people. And, like, the other cool thing I guess, last last thing I wanna show.
Speaker 4: Like, this is your this is your agent's homepage. And, if I refresh it, because we just we just Lloyd this thing in the demo, there's the apps that it's Consulting. And so this is I like a homepage for your AI agent. Right? And, yeah, you could so, like, your friends could look at, like, hey, what's, you know, what's Seb been you know, Seb's agent been deploying lately and, like, here's an app.
Speaker 0: Are you
Speaker 5: familiar with the AirDog now? No. Okay. That it's something similar.
Speaker 4: Uh-huh.
Speaker 5: So I'm wondering, like, if I wanted is so I can go to your GitHub and get that skill.
Speaker 4: Yeah.
Speaker 5: And then then Two what? No.
Speaker 4: That's all you need. So, like, your agent will take it from there. Right? I can show you what the skill is. She just pointed me that that we're out of time.
Speaker 4: So we can maybe talk after, but, like the skill is it has a script to do the install, and then it tells you how to use the it tells the agent how to use the thing. So I, it's gonna run, basically, the it's gonna run, here. Exposed. Two to Lloyd exposed name port.
Speaker 5: It can host my own apps
Speaker 2: Yeah.
Speaker 8: On my
Speaker 4: own On your own hardware.
Speaker 5: And,
Speaker 4: And you can give and you can give you the URL to, like, your grandmother and she and it'll work for her. Right? She doesn't have to install anything on her.
Speaker 5: Okay. Cool. Thanks. Thank you.
Speaker 3: Not to nerd smack you, but, what what what if I had, like, I opened my laptop and had a a MCP server Mhmm. Running on that. I could, like, maybe serve that MCP server to the agent that's Two over in your house. I
Speaker 4: don't know. Yeah. So so there's, like, agent to agent stuff that you might wanna do. So, like, that's not through a web browser either. And so the so the other thing that I do and, like, I didn't wanna get into this because we don't have a ton of time is that, like, if you install the agent, it will, it will also Architect, like, it'll also intercept DNS to other agents.
Speaker 4: So you can do, like, direct, like, direct NAT punch through to to other to other folks with p 2 claw installed. And so you can do back end API stuff, maybe agent to agent I stuff. Like so it's kind of a primitive, I think, for, like, agents to communicate with each other, At I, like, from like a infrastructure everyone's has has these Mac Minis. That's kinda that's kinda how I'm thinking. I'm thinking I, what are primitives?
Speaker 4: And so 1 is like, I wanna host websites. Another is I, I wanna be able to talk to other agents. Maybe another one's email. Right? Like, maybe I'll make email inboxes for each of these guys.
Speaker 4: I don't know. We'll see. But, yeah, that's where I'm at. It has a website here actually. It's just Two.
Speaker 4: There you go.
Speaker 2: You wanna go straight into the voting game?
Speaker 0: Yeah. I'll make a special. Alright. So he didn't present, but Ron created something that we're gonna use now. You're Hercules.
Speaker 0: So, if you've been to our tinkerers before, you New we used to have the whiteboard. His name's on there and counting votes. Ashley remembers because he's been around long enough. Now we got something better that Ron created for us
Speaker 4: for voting.
Speaker 2: Over the voting page. And this page was actually made by. And it's running on Intel star on a note from my house.
Links
Tech stack
Finding related talks...
Compose Email
Sending...
Email preview
Loading recent emails...