What if the AI arms race has no finish line?
In this episode of Your AI Injection, Deep Dhillon sits down with Rob May, CEO of Neurometric, to challenge one of Silicon Valley's favorite assumptions, that whoever builds superintelligence first gets to keep it. Drawing on over a decade of AI investing and two prior startups, Rob lays out a public prediction he's willing to stand behind. He believes that once any lab crosses the threshold into true superintelligence, open source versions will follow within months, not years. The two also get into the "jagged frontier" of AI capability, the strange reality that today's most powerful models can be brilliant at one task and shockingly bad at the next, and what that means for how companies actually deploy AI. There's also a simpler question underneath it all. If AI keeps getting cheaper and easier to use, will we just find more excuses to use it, whether we need it or not?
Learn more about Rob May here: https://www.lin kedin.com/in/robmay/
and Neurometric here: https://www.neurometric.ai/
Check out some of our related content here:
Get Your AI Injection On The Go:
xyonix partners
At Xyonix, we empower consultancies to deliver powerful AI solutions without the heavy lifting of building an in-house team, infusing your proposals with high-impact, transformative ideas. Learn more about our Partner Program, the ultimate way to ignite new client excitement and drive lasting growth.
[Automated Transcript]
Rob: I think there's this belief that like whoever gets to AGI first, even if it's just a matter of hours, they just, the thing self improves and it takes off and you can never catch it. I am willing to bet, and I'll say this publicly, it'll be the exact opposite. Whenever it happens, within three months, we will all have access to open source super intelligence.
Rob: And I don't know when that'll happen, but whenever we hit that threshold that people agree it's a thing, I think it'll be available to all of us. I, I don't know that anybody's gonna win
CHECK OUT SOME OF OUR POPULAR PODCAST EPISODES:
Deep: Hello, I'm Deep Dhillon, your host, and today on your AI injection.
Deep: I'm joined by Rob May, co-founder and CEO of Neuro Metric. Rob earned degree's in electrical engineering and an MBA from the University of Kentucky and has spent his career building and investing in AI companies at Neuro Metric he's working on the next layer of the AI stack, how teams empirically select, combine and route models at inference time as cost, reliability and vendor independence become first order concerns.
Deep: Rob, thank you so much for coming on the show.
Rob: Yeah, thanks for having me.
Xyonix customers:
Deep: Usually I start off by [00:01:00] having you walk us through your solution stuff, but I'm kind of want to do something different and jump into your investing background and how that like landed you in, the role that you're at.
Deep: So talk to us a little bit about your, we've put your VC hat on again. Tell us a, uh, briefly about your VC background and give us a little landscape perspective that sort of caused you to really dig in on NeuroMetric.
Rob: Yeah, it's a great question. I started investing in 2015.
Rob: I, and so almost, 11 years ago, uh, I wrote my first check into an AI startup, and that was because. In, um, late 2014, I had sold my first startup, uh, had a really good exit, decided to start doing some angel investing, wrote a small $15,000 check into this, you know, uh, machine vision startup. And I was thinking about what I was gonna do next.
Rob: And in 2015, if you looked around, slack was kind of taking off. So it was like corporate messaging. Everybody's talking about IOT, everything's gonna be connected to the internet blockchain. I read all the papers, didn't get it. And then AI, Google had come out with their word vectors paper the year [00:02:00] before that had the famous like, king, queen, uh, yeah, yeah.
Rob: Analogy. Yeah. If you've seen that one. So I decided like, oh, AI is gonna be my next thing. So I started doing a lot of ai, early stage angel investing, and that attracted the interest of a firm called PJC in Boston. They invited me in as a partner and a general partner to focus on ai. And I did that for a couple years.
Rob: And, I loved. Many parts of the VC job, like talking to people who are trying to build the future and hearing all the wild ideas and everything else was awesome. I did not like any of the pieces about managing a partnership, working in a partnership or running a fund. So, one of the things you have to do is to, to be a good investor is you have to be like, Hey, maybe you should think about this.
Rob: You can't be. And you know, as a, as I was a two time CEO at that point, and I was like, can I just sit on a sales call and tell you what you're doing wrong? Because,
Deep: well, you can do that. So, I mean, like, once you earn the trust of the founders, you know, you, they, I don't know, they can be very entertained by that idea.
Rob: Yeah. I don't know. I, I, I was a little too control focused, so I came back to the [00:03:00] operating side. I'll tell you the biggest thing that influenced my interest in neuro metric from my investment background is the thing that I learned as an investor and also as an entrepreneur now, is when you cozy up to really big markets and really big problems, you can do many things wrong and still be successful.
Rob: And what I love about what we're working on in neuro metric is we play in the inference market. And inference is probably going to be a multi-trillion dollar market over the next 6, 8, 10 years. It will be one of the largest markets in the history of the world. As all these work tasks move into inference.
Rob: And what that means is if you pull off a tiny sliver and you optimize for an edge case on the inference side, that could be a billion dollar outcome.
Rob: It could be a big company.
Deep: Yeah. I mean, I think that when I, when I saw your. A description, and we're gonna get into it in a second. I mean, I was thinking of something like PagerDuty.
Deep: 'cause the first time I saw PagerDuty, I was like, this is so, this is so narrow. Yeah. And so obvious. But you know, here they are this like massive company and it's super surprising. But, um, so yeah, I mean, with that, I [00:04:00] wanna come back to your investment perspective later on in the show, once we've kind of dug in on your business a little bit.
Deep: First outta the gate, I'm gonna say kudos for diving in and getting your skin in the game. I'd say the vast majority of, you know, VCs don't do that. Like, once they get out, it's very easy to just sort of, you get this nice perch, you can sort of see all the stuff and you kind of, it's very easy to get addicted to this feeling of being the sossay and like for, you know, and predicting where to put stuff.
Deep: Yeah. But at the end of the day, I mean, you're, you're, you're pretty much a banker, you know, or like, you're, you're a financier and there's a lot more fire in the, in the hot seat of actually running a company. So, you know, kudos on that. Yeah.
Rob: And actually, yeah. Well, and actually just the thing I'll interject there is, um, and VCs like to think about themselves a little bit as like the mavericks of capitalism that are taking all the risks.
Rob: But part of what I didn't like about the industry is actually I, and there are some, there are some who are great, but by and large it's a liming based industry where everybody follows the handful of really novel leaders and wants to do the similar things. And so it's, you know, a lot of times you couldn't invest, I couldn't invest in the [00:05:00] stuff that I really thought had potential I wanted to invest in.
Rob: And that was,
Deep: And depending on the stage you're investing in, it can appear as if you're taking risks, but in reality, if you're running like an early stage fund, you know you're going through 30, 40 companies, you're expecting like a really large failure rate. I mean, you could close your eyes and put your finger in the wind and just grab stuff and if you're in the right group of people with the right basic credentials and you're not.
Deep: Way off on your own. Yeah. I mean, you might hit not horrible returns that with that approach. So yeah, I, I don't think that makes you a genius who really understands, you know, the space. But on the other side of the fence, I have worked with investors who are really, really sharp and really have a unique perspective because they see so many companies at the stage that your company's in, and they can just like, give you this really valuable advice that can be really timely, and can really change the direction of the company.
Deep: But
Rob: yeah, I, I remember I, at my first [00:06:00] company, so my first company did backup for cloud competing applications like Google Apps and Salesforce and Office 365. There were four players in the market. We were the largest. Most VCs I talked to, some of 'em didn't even understand what we did. They were like, you know, how do you compete with like Carbonite?
Rob: And I was like, we don't compete with Carbonite. And, and then most of the good VCs knew. Us and the number two player, when I went into pitch Sequoia for the first time, which they did not invest, they knew the fourth player in the market, this small company, whatever, and they're like, how often do you see these guys in deals?
Rob: And what do you think about their feature that does this? And I, and, and I literally was like, how do you know about them? Like, nobody, nobody even knows they exist. We've, we've only seen 'em in two deals ever. Um, and so it does show you like, I mean, those guys like worked hard and went deep and so I, there's definitely, yeah, we're getting off topic now, but like, yeah, there were definitely differences in the, in the quality of the
Rob: VCs.
Deep: Yeah. I remember I pitched at Sequoia like a long time ago, 15, 20 years ago. And. Uh, yeah, I got, I got just torn apart. Yeah. Because I was like a lone tech guy running [00:07:00] around, like trying to like sell, sell my, my company. that was not yet even spun out of another company and it was like a complicated deal.
Deep: And anyway, it was like, it was way before I knew anything about business, how the business world worked. And so it was entertaining. It ended up being fruitful in the end, like I ended up getting introduced to the right people and got an exit. But anyway, let's dive in. So on neuro metric, what's the problem that people have without neuro metric solution?
Deep: And for the sake of our audience members, maybe a few that aren't as familiar with what it even means on inference. Maybe like back up a few steps and describe, you know, what kind of happens at the API typically without neuro metric and what people end up doing. And then why, why you think centralizing that point Makes a lot of sense.
Rob: So, so, so we look at the world through the lens of how AI matures your company, because what we have seen is AI maturity, which is how far are you on your AI journey? Uh, doesn't [00:08:00] really, like, it's not small companies, it's not big companies. It's like, kinda like your, it's just kinda like your attitude towards ai.
Rob: How early do you adopt it? How experimental are you? How much are you try and stuff? We solve the problem for companies that are past that first part of the journey and move into the optimization phase. And so to give you an example, what most companies do is they start kicking around proof of concepts on either open AI or Anthropic or Gemini, right?
Rob: Like one of the top, top three models, uh, which makes sense. That's how you should start it, right? You're exploring capabilities. Can I build an agent or some thing that, you know, does the thing that I need? The problem is, and, and then what they do next is they, they go to production with something finally, after.
Rob: 6, 12, 18 months, they go to production with something and then they go, wow, what if the Gemini API's down? What if the Claude API's down, we should fail over to something else. So then they have two frontier models in production, and that lasts for a while until they get traction. They go, wow, we're spending $50,000 a month on inference.
Rob: Right? and so, you know, for people that don't know, a lot of the last few years have been about training models. You train a bigger model and you do all [00:09:00] this data, and that's when you hear about it cost a hundred million dollars to train this model. It's like they just, they were teaching it and teaching it and teaching it and having it go through the data and fix itself until it got all the data right.
Rob: And then inferences when like, the model's trained and now you're just running something through it. Right. Um, and saying like, you know, it's, it's like, Hey, hey, I've got this glass here. And it's like, Hey, is this a a, a water glass or a wine glass? And you're running it through and it's telling you the answer, right?
Rob: 'cause you've trained on all these water and wine glasses. The reason inference is expensive is because as these models get bigger. You have to store the layers of the nodes in memory, and those have to be shifted in and out of the compute parts of the GPU. And the bigger the model, the more times that has to be shifted in and out for every time you run through it.
Rob: So it gets really, it happens really fast, but it's still the, these models are big by computer program standards. So like if you think about an operating system, like if you're a person watching this and you're old enough to remember when they would send you a stack of discs to install the Windows operating system, right?
Rob: And you're like, oh man, it can't even fit on one [00:10:00] disc Windows might be a 12 or 15 gig program, right? Like an operating system. I mean, these models are 7, 8, 900 gigs, right? You, you're like really big models. So they're, they're hard to run even on one GPU and they, and they take some time. So that's why inference is so challenging to make fast and that's why Claude and OpenAI and Gemini sort of speak to you and stream stuff out and whatever, rather than just flash up an answer.
Rob: Um,
Deep: yeah. And not to mention there's sort of like an overwhelming array of options too, right?
Rob: Yeah.
Deep: So there's lots of different models that modeling landscape is gonna get more complex, not less.
Rob: Yeah.
Deep: I, I think, you know, like you've got Kodak like programming optimized models. You've got really large kind of overall reasoning models.
Deep: You've got models that are starting to just sort of excel at writing and, and authoring. And then you have the whole world of private models where you can run something maybe all the way down on your laptop and up, and they're all ultimately very, very similar. [00:11:00] But being able to kind of dynamically make those decisions about like, if you're just trying to summarize some text, you do not need, like, you know, A-A-G-P-T five or a Gemini three or whatever, right.
Deep: You know, like you can go with,
Rob: and that's what we see, right? We see companies that are going like, wow, we're spending so much on inference. Do all of our workloads really need to go to this? Like, look, you're always gonna need frontier models, right? At least for the next 10, 15, 20 years. You're gonna need people pushing the boundaries.
Rob: The models can do a lot of stuff. And particularly 'cause they can handle lots of one-off tasks, which in work you have a lot of one-off tasks. But when you have things that you do repeatedly, like to your point, like yeah, can you, can you summarize this blog post? Like you don't need GPT five for that?
Deep: No.
Rob: Um, and in fact if you're doing it a hundred times a day, you can probably build a small language model that'll do it better than GPT five for like one 20th of the cost and 10 times faster. Because again, the memory thing, small models run faster 'cause you load less in and outta memory. Right. Um, and so what we do is if you're like, well I would like to move off GPT [00:12:00] five, but I don't know what models I should try.
Rob: And there's, you know, 200,000 models out there now on hugging face and variations. It's like, well we can take your workloads and we have a giant, you can think of it as like a giant test harness that runs a bunch of stuff and we. We add, and maybe we'll talk about this later in the podcast, but we add test time, scaling algorithms on top.
Rob: We try some different stuff. So we try different models, we try different test time scaling algorithms to probe those models. And we come back and we say on your data, here are the things that work best for speed. Here are the things that work best for, um, cost, uh, at sort of similar levels of accuracy.
Rob: And if you want better accuracy, we can also tell you that.
Deep: So, so I'm gonna let, let me try to like, summarize what we've sort of talked about so far and kind of make it maybe a little explicit and maybe, take a particular example. But I just wanna make sure that kind of everybody drops what we're talking about.
Deep: So let's, let's take something really simple like here's blob text, summarize it. So typically, there's, there's the point in your, that's an example of like a stateless [00:13:00] transaction. You've got the full blob of text, it's all ready to go. You pass it into like an open AI or Gemini or whoever.
Deep: Um, you initialize, the LLM endpoint somehow, typically. And then there's like an API endpoint where you, you're passing that text plus your prompt. And your prompt is like telling you what to do. So if everybody's used, you know, the LLMs themselves. So you know how, you do that via chat, typically.
Deep: You're just like talking to the thing. Well, in the API world, you're, you're sort of defining that and it goes up with the, with the text and bam, that comes your answer. So my understanding of what you guys are doing is you're like, here's our endpoint. Our endpoint, is aware of all these D we have our own API keys and everything for all these other models.
Deep: And even within each universe, Gemini's universe, Anthropic's universe, open AI's universe, there's this other universe of all these detailed models. And we'll take your call and we'll pass it through all those places and we'll expose it to an [00:14:00] efficacy test of some sort, which I'm presumably me as the caller is defining.
Deep: So I'm sort of giving you a thousand or something examples of blobs, of text, examples of optimal output. And I, I'm assuming you, or they are deciding some kind of metric to sort of say how far away from the optimal answer is this model and that. And in that sense, you can basically now start meandering about hitting all these different APIs and come back to them and say, Hey, it's subject to your ground truth and your efficacy test.
Deep: Like this is. This is like the latency view, like how fast it or long it takes to get Yeah. And this is the efficacy view. And then you could probably come back with some kind of intersection and like, we think you should sit somewhere between here and this is the cost view.
Rob: Yeah. Yeah. That's a great, that's a great explanation.
Rob: And, and what's interesting about this business, we actually tend to see the lat latency as the driver more than cost at this point in the, in, in the market. And I'll tell you why. I think, I think as people are [00:15:00] building agents and you design your agent to make, say, say it's a 12 step agent and you make one call each time, well, you know, you look at the, the Anthropic APIs, they might be 800 milliseconds to two and a half seconds per call.
Rob: Well, at 12 seconds you might have a 25 second agentic workflow where you're like, well, if I could pull nine of those 12 steps off and run 'em on a 200 millisecond or less small language model, that would accomplish the same task. Like if I don't need clawed. For all 12 of those steps in the task. Just for nine of 'em or just for three of 'em.
Rob: Right. And the other nine can be put to some, maybe you could take that 25 second a agentic workflow and make it a three second Ag agentic workflow, which is huge for your user base. Right. to give you a crisp example of like something we did with a customer that shows the use case, we had a cu one of our first customers came to us last year and was like, Hey, um, we ingest thousands of pages of documents about companies and we extract like they're e-commerce companies and we extract like entities and uh, you know, lots of [00:16:00] information and we build a knowledge graph and blah, blah blah.
Rob: And we go, okay. Uh, and they're using, I dunno, GPT for one of the variants there. And it was taking like a day and a half to build that knowledge graph. And so we ran out a bunch of models and we said, Hey, we, we actually found a model Lama for Maverick that can run this at. Four times faster and in one 10th the cost.
Rob: And they go, oh, well the cost savings is great, but the four times faster is actually bigger for us because now we can do more stuff and make more decisions if that's happening faster. So those are the ways that we help people out. And then as you can imagine, as more models come out and this becomes more dynamic, um, you know,
but
Deep: I'm wondering if, like, if, if you're, if you're making a decision based on speed and cost and ignoring efficacy, then maybe you're just making a, like leaning on your agent to go do a bajillion calls to make up for it somehow.
Deep: So it seems like you need all three of those guys to make
Rob: any Yeah, no, no. I, yeah, I should clarify. That's a great point. We, we do look at, at [00:17:00] efficacy, and there's two ways to evaluate it, right? One is sometimes people have really good eval sets. They're like, Hey, when you give this prompt, this is what should come back.
Rob: And we can take those eval sets and we can test it, and we can say, well, this, this model gives you the results you want faster. You would be surprised how many people don't have good eval sets and we'll say something
Deep: like, I would not be surprised at all because, you know, I run an AI consulting firm, Xyonix, and I cannot tell you how many times we like begged, plead to get people to like step up and define the ground truth test.
Deep: And we, and or even just grant us time to do it. And in the old world before all these big models, it was easy because I could just say, look, we have to have the ground truth or I can't give you your model. But now that they can get a model without a ground truth, they'll weasel out of it nine times outta 10.
Deep: And they, because it's, it's like harder to explain how to reason about uncertainty not only with, you know, non-engineers but also with engineers. 'cause very, very few engineers are actually trained probabilistically and [00:18:00] so they don't really get it. And, but you can usually explain it to an engineer.
Deep: But for non-engineers is like really challenging to communicate.
Rob: But you, you and I are double E, so we had to have all that signal processing crap, and we had to learn that probabilistically. That's how you have to think about the world. Um, yeah. Do you, are you surprised at the number of times you see people where you have some use case where they're like, we're taking all these 2000 pages of industry reports and you know, we, they, they're about, you know, the energy industry and we want the, uh, we're asking g PT five to pull out all the stock symbols, and you're like, great.
Rob: And it gives you a list back of stock symbols, and you're like, stock tickers that are in there, and you're like, uh, and how do you check to know that they're right? And they're like, we don't, we just assume that it got all the stock tickers. And I'm like, well, you know, they're like, it's crazy.
Deep: I mean, you're, you're preaching to the choir because this is, it's so hard in the modern era with these LLMs because people get the illusion of quality really fast.
Deep: Right. [00:19:00] And I can usually get engineers to buy off because I just put a very simple. scenario to them. I'm like, okay, well here's what's gonna happen if we don't do this. Some exec at your company is gonna have their favorite thing fail, and they're gonna breathe down your necks and drive you crazy.
Rob: Yeah.
Deep: Until you address that thing. Now, when that happens, you can either like one off their thing every five minutes and they just get more and more pissed off every time it keeps failing. Or you can point them to this like statistically valid set that you've got and you can explain like why, how you measure things the way you do.
Deep: And in exchange you can tell them, I'm not gonna fix your problem, and they'll have no choice but to accept it. But you can say, but I will include your example in my data set. And so engineers will get it really fast, but like non-engineers, they'll get it or say they do or nod, but they don't want to give you a time and resources to do it.
Deep: So I'm curious, how do you deal with that? Because I, I, it seems like [00:20:00] your. Limited by this problem. And, and like, you need to like, make it really easy for these folks to like, define their ground truth sets. And I'm curious how far down that rabbit hole you guys go, 'cause you're in a good position to ha to basically build a ground truth labeling tool as well because they're already uploading, like to just stick with that old task.
Deep: They've already got the reference text, they've, and you generated the response and you can give them an end point that says, this sucks, or this is great and there's your ground truth, right? So
Rob: yeah. It's, um, so, so there's, there's a couple ways to deal with it. Uh, on, on, on the one hand, uh, part of it's an education problem.
Rob: So some of the stuff that you mentioned, we try to explain to people, like, these things are probabilistic and it, and at some level it starts to get hard. Like, how do you evaluate a human on this task, right? It's like, you wrote a great analyst report. I'm like, this looks like a good report to me. Could it have been better?
Rob: Maybe I, I don't know. Um, the way you get there is like doing lots of reports, right? As, as a human. So, so we're, we're gonna run up against this problem as these models get better and better, the ambiguity that we have with, with. With [00:21:00] humans sometimes. What, what we typically tell people, I mean, oh, I love when people have a good eval set.
Rob: Right? Uh, that that's the best. We try to push people to build one. We have some partnerships with some data labeling companies that can help. And, um, that's ideal. When they don't have that, we typically try to do a sort of relative evaluation. So if you go back to my example of like, Hey, I, I gave you a couple thousand pages of text about the energy industry and you gave me some stock symbols back, we'll just run that on some other models and we'll be like, well this gave you like 99% the same stock symbols back as the other thing.
Rob: We don't know if those were right. 'cause you never told us if those were right. You don't know if GPT five gave 'em to you. Right. But we mimicked what GBT five did. Like this model gives you the same answers. GBT five did so and it's cheaper or faster. I see. So you benchmark we these relative valves
Deep: relative to the beefy model.
Rob: Yeah.
Deep: And we also, that's kind of what for tech stuff, I mean that's what most people are doing anyway to generate their ground truth. And for straightforward tasks that works really well. But [00:22:00] for subjective stuff, like even something like summarize this, um, you, it's not obvious like the, you know, how to actually say whether or not the thing the model gave you is better than the thing that you're human actually said, or that you're human, like statistically validated against a higher stack model, right?
Deep: Like it's, you have to use, typically you have to use an LLM for that too. Like in the old days, you know, we had very specific mathematical metrics that we actually understood, but now people like those just get outperformed so quickly by, by the LLMs that we're like building LLM comparisons and now to communicate that gets more and more like, you know, complicated.
Rob: Yeah, yeah. No, that's a great point. Um, and there's, you know, there's, there's a lot of things you can do to try to try to make these tools better. And a lot of it is like handholding people and explaining to them like. He helping them understand like, here's how you should think about it. , Yeah, sometimes all we can do is sort of the, the relative [00:23:00] comparison.
Rob: And now, now there, there are use cases, so like take a product recommendation use case, right? Which is like, Hey, I'm gonna take a description of this product and I'm gonna recommend five similar products on the webpage. That's a use case sometimes where we can get away with like act with actually saying like, Hey, this model's twice as fast and half the price.
Rob: It's not quite as good. It's like 85% as good. And that's a use case where you might go, oh, like that's worth it because of the five products I show you. Like, as long as you get the best one and you buy, like there, there are AI use cases like recommendation systems where maybe you can even take a little bit of, quality decline if you get a big enough cost improvement.
Deep: Absolutely.
Rob: And there's other places where you can't.
Deep: So, so just to be clear, do, do you actually build in the feedback mechanisms so that they can define their ground truth, like piece by piece so they can go ahead and run and you get, you get them their cost and their latency numbers, but when they see mistakes they can flag them easily either via API or US and, and therefore they can start to build out their ground truth.
Deep: 'cause it seems like that's a [00:24:00] fundamental part of the conversation is like, Hey, we think you're doing great. We also only have eight entries in our ground truth. So it's kind of a bullshit statement that we're making. It seems like building that in is kind of key.
Rob: There's definitely ways where you can go back.
Rob: There's definitely ways where you can go back and categorize like, 'cause a lot of we try to do is break that. 'cause we're, we're trying to provide the best model for your AI workflow. So maybe you have customer support responses, maybe you have tech summarization, maybe you have recommendation engine, maybe you have some coding or whatever.
Rob: So we're trying to say like, oh, for this, for your coding model, use this for your recommendation engine, use this, blah, blah blah. And so we do allow people to like recategorize some of those and be like, oh, that prompt wasn't that, or whatever. Um, when they don't have a hard eval, there's not a way to say like, that prompt and that answer was wrong yet that, that is coming.
Rob: That is a product feature that has been requested is coming. Part of the way we get around that is we've added functionality like, um, we can ingest l use data. So if you have some kind of um, LLM [00:25:00] observability platform that's monitoring and watching things for you. We're trying to add the functionality to ingest all that data, and then we can take that data, you know, oh, uh, your, your, your length use instance saw you call this LLM and get this back.
Rob: And we can use those input and output prompts as test data. We can run that on a bunch of models and say like, well, this model will give you the same things back, but faster. So there, there's, there's little tricks we can do and we're always trying, trying to figure out, like little hacks that that'll tell you things about the AI world, you know, without having to test everything.
Deep: I like the idea of a two-pronged approach. One is benchmarking relative to the best foundation models. Um, that makes sense. 'cause they don't actually have an alternative, so it's not Yeah. Right. Like other than breaking down their problem and because actually that's probably the alternative is like, if they're not measuring efficacy, like end to end for a task that hits like 50 LLM calls, um, you know, [00:26:00] but if they, if they are then, and, and you're still weak, then you could start to like inter, I mean you should anyway have observability into like different parts of the stack, like wherever you're hitting, hitting your classifiers or your LLM.
Deep: But I also like the idea of like giving them endpoints and along with user roles so that they can directly provide actual authenticated human feedback either coming throughout through the product or, you know, from their internal people or whoever.
Rob: Yeah. And the other thing we do sometimes we just, we just allow manual feedback.
Rob: So we just say, 'cause sometimes people just wanna run a handful of models to see what they do. And so it's like, well, hey, great, you know, you gave us a hundred prompts. We ran 'em on these five models, we ran, you know, two, we ran chain of thought and best of end on each one of 'em. And, uh, you know, here's, you know, 500 responses back.
Rob: Feel free to look through 'em and see what worked for you. Right. We, we can also do that sometimes.
Deep: The, the other question I had for you is, running these tests is really time consuming and expensive. So how do you guys think about that and like, [00:27:00] you know, especially like to get a statistically meaningful set, you know, you need to be into the hundreds or thousands or tens of thousands that takes Yeah, a lot of time and money, especially if you're trying to benchmark them relative to a foundation model.
Deep: So what are the kind of like knobs and dials that you give the end user and how, how do you get them to think about it and then. How do you guys go back and sort of optimize simply the text test, test execution path itself?
Rob: So we have some of our own data because we run a leaderboard where we test a lot of models with a lot of test time, scaling algorithms, uh, on specific work task data sets, not on a lot of the traditional academic benchmarks.
Rob: So sometimes we can use that as a comparison. Some of the data we have from other things. And as we get a broader customer base, we can start to understand, um, you know, this matches up. What's the statistical significance of seeing? We've seen this model a lot. Like you learn a lot about these models. Like, to give you an example, uh, we were working with the CRM Arena data set, [00:28:00] which is a series of sales tasks that came outta salesforce.com and we ran these different models on each task and we showed that models perform very differently.
Rob: No model wins on every task, right? Um, no model even wins on half the tasks. So then we said, well, could we train a small language model to do this? And it turns out, yes. We trained Quinn three, 4 billion parameter model to beat frontier models, all the frontier models on this one task of lead routing with the Salesforce CR Marina dataset.
Rob: And that took a fair bit of work. But the interesting thing was models of a similar size to that Quinn three, four B parameter model, we could not ever get there. So there's a thing of like, well, why could this Quinn model do it? But some of the other smaller models of similar size couldn't be trained to do it.
Rob: Like these, these models are very stochastic in how they've been trained, in what they know and what they can do. And so part of what we're getting really good at is being able to say, like, understand the statistical distributions of these models because we probe them so much. And so [00:29:00] we've, we actually have a unique search algorithm that we're patenting, which is, um, ha have you heard this term, jagged Frontier about ai?
Deep: No.
Rob: So it, it came out of a, a, a Harvard paper from, from two years ago. Where, um, they're basically like, uh, so the problem with AI is, uh, there's this most technology sort of advances at this frontier of like, you know, uh, uh, stuff. But with AI, they're like, you launch a new model and there's some things that it's way better at than you would expect it, and some things that it's way worse at than you would expect.
Rob: So it's like this jagged, sharp frontier of peaks and trs. So we came up with a new way to search that because what you're really trying to figure out is, okay, if I have, you know, if I have five models, I could just run it on all of 'em. But if I have 500 models, how can I pick the 20 that I might actually want to test that might be able to do this?
Rob: Um, and you can use a lot of dynamic adjustments to do that. So the, the example I would give you is like if you're, if you're a high jumper and I set the bar at four feet and you clear it by a foot, I don't need to set it to four foot [00:30:00] one and see if you can still clear it. Like I can move the bar up quite a bit.
Rob: Right. These dynamic adjustments. Mm-hmm. So a lot of our system is defined to like. How do we look at all the data we have? How do we make dynamic adjustments to be like, oh wow, that performs surprisingly well. We should do more over there.
Deep: So one, one thing I didn't quite understand is you mentioned that you guys have your own ground truth regardless that you apply across the foundation model.
Deep: So it's not like a customer centric one. It's like your generic, test. What's your driver there? Like, is your main goal to get kind of the lay of the land of the models or do you publish that and sort of use that as a lead gen opportunity for your clients? Because it seems to me like a Yeah, yeah,
Rob: yeah. We published a leaderboard is that these models vary on a per task basis so you can see which ones are right for the tasks closest to your use case. And then, we showed that these test times scaling algorithms like chain of thought.
Rob: Best event or beam search or whatever,
Rob: but you might have a task where chain of thought and approach, uh, [00:31:00] wins on this model over here. And so when you think about test time, computer test time scaling, a lot of it is for us to test that out so that we understand it and also for legion purposes for people to po 'cause the kinds of people that are poking around and interested in that are probably good customers for us.
Deep: I mean, honestly, so this idea came up, I was chatting with somebody like three or four years ago, I think it was like after GBT two came out. And the idea was like, these guys are all gonna, they're either gonna lie about their stuff or, which is unlikely, but they're gonna like sort of. Broadcast the foundation model.
Deep: Companies are gonna like really publish and publicize the stuff that they're really good at and underemphasize the stuff they're terrible at. So like none of them, for example, to this point, has gone and published how, at least last I checked, how well they do on assessing whether or not somebody's a paranoid schizophrenic and they're playing along with the delusions.
Deep: 'cause they all have still terrible
Rob: Yeah.
Deep: Data on that, right? So there's, there's room for from a non-biased or just like [00:32:00] orthogonally biased perspective, you know, like in your case I'd say you're probably just biased in a different way. You're biased towards what your customers care about probably.
Deep: Um
Rob: Right.
Deep: But it's, it's still something that makes a lot of sense for outsiders to like a look at and, and be like, oh, I haven't thought of some of those use cases. And it seems like a great way to get customers too. I'm, you know, they can sort of play around and sort of, do you have a tool that like, helps them find the nearest data set that you guys have numbers for or like
Rob: We are, um, yeah, so we, we, we have a tool called Explorer and what it allows you to do is test and compare models on a prompt by prompt basis.
Rob: So you wouldn't wanna use it for any kind of mass analysis, but if you're like, huh, I'm doing prompts that are of this nature and I just wanna try a couple different models and see, you can select your models, you can set the prompts like Yeah. That's something that we offer and that, it's not as popular as people just looking through the leaderboard, but some people do go use it.
Deep: So one of the things that I've found challenging is the, [00:33:00] is like when you're trying to assess how well these models work on a particular task, And even if you're trying to assess cost and latency, the challenge with these models is they're so non-deterministic. And so you can make the same call two or three or four or 10 or 20 times.
Deep: But at the same time, it's very expensive to just keep calling them expensive by latency and, and by dollars. So I'm curious, like, how do you guys think about that? Like how, how many times do you run the same call to even understand the drift and be able to figure out, a statistically meaningful metric on how you perform?
Deep: Like, going back to my example, you got some text and you're gonna summarize it. You, you gotta call a model technically like a, at least a handful of times to figure out how far the different permutations are from it.
Rob: Yeah.
Deep: How do you guys think about that?
Rob: Yeah, it's, uh, so we don't, we don't publish this data.
Rob: We did one blog post on this topic. We will eventually publish this data in the app when, when we wrap our heads around a little better. But, um, yeah, so we typically do for any model. [00:34:00] Platform algorithm combination, we run it five times. So to give you an example, if you're testing lead qualification task, uh, from Sierra Marina and you're running it on Amazon bedrock Llama three 70 B with best of in weighted, uh, average, we would run that combination five times.
Rob: And one of the things that we found is, uh, these models, depending on the model and the platform that you're running it on, they actually vary quite a bit between calls. Some of them do on a latency or cost basis. So we give, so we're gonna give you that spectrum. So the, because, because people don't engineer this way today, but you can imagine of like, look, we have tight latency criteria and so if the, you know, we, it needs to be, between 60 and a hundred millisecond response time, we can tell you like, oh, this model always is in between that response time and other, this other model is.
Rob: Performs better, but is maybe like, you know, sometimes it'll be 50 milliseconds, sometimes it'll be 120 milliseconds. Like can you, [00:35:00] like, if you can't tolerate that, you can't move to that model. Those kinds of engineering sometimes
Deep: 8 38 seconds, you know? So.
Rob: Yeah. Yeah. And so and so then the question is like, how consistent are these across performance runs?
Rob: That's data that we're collecting. And we'll be able to tell you soon as well, because I think to your point, I think people are gonna wanna know that.
Deep: Yeah. I think looking at the standard deviations across the latency makes a lot of sense because. Uh, a actually across all of these things, across the cost latency and the efficacy.
Deep: Yeah.
Rob: Because
Deep: they just move around and, and then there's like other stuff that we haven't even talked about. I mean, I'm sure you've noticed that there's times when open AI is deploying a model or something where things just start falling apart and it doesn't Yeah. It doesn't mean that your call blows up.
Deep: It just means it's like the latency blows up or the efficacy plummets, like weird stuff happens because, you know, like someone's playing around internally with like how deep the models go, you know, and Yeah, like there's, there's like a ton of stuff determining how, how [00:36:00] much reasoning gets applied to your task, you know?
Deep: So.
Rob: Yeah. And oh yeah. There's so, so many variables here, right? Like, like when you're, when you're looking to, uh, we have seen a lot of that. We've even seen that if you take an open source model like the Quinn or Lama or something, and you run it on a different platform together, AI versus GCP versus Amazon Bedrock, the accuracy might vary for the same prompt.
Deep: Oh
Rob: yeah. So you, because you, you know, each of those platforms is thinking about how to best optimize those things for their customers. So, you know, the, the way they manage, like KV cash in the GPU might be different or something low level like that. And, you know, I, I don't know what it is for that platform or whatever, but you can see where those things could have an impact.
Rob: You know, and it's complicated too when you're designing these systems because you're like, let's say your whole cost is like, well, we would like to drop our LLM cost by 10% this year. Well, there's a lot of way to do that. You could probably rewrite your prompts that might do it.
Rob: You probably move to a cheaper hosting platform. You probably move to a open source model or a cheaper model, or at least part of your calls to a cheaper [00:37:00] model. Like, you know, you could probably, uh, there's, there's probably a whole bunch of things you can,
Deep: you can host your own models, you know?
Rob: Yeah, yeah, exactly.
Deep: There's all kinds of stuff you can do. So one thing that I saw that I thought was interesting from your platform, and I wanted you to talk a little bit about it, is, so, and tell me if I got this wrong, but when you make an LLM call, sometimes they fail. Sometimes you, you, you know, like developers will, sometimes you're, you're making backups to the backup to the back sometimes.
Rob: Yeah.
Deep: Uh, you know, sometimes. So there's like, there's like a lot of different things that can happen. Um, sometimes you categorize something and based on the categories results, you might go one way or the other way. And I noticed that you have some kind of, you know, directed graph like approach where you're pulling in all, you're, you're trying to like pull in that logic into your world.
Deep: So that, which makes sense. 'cause if you're trying to drive these metrics, you, you want a lot of that sort of in there. So can you speak to that a little bit? Like how you guys think about that and, and, and maybe what's the [00:38:00] state before and after using your product? Because before I'd ever really thought about your product, you know, we're sort of managing that in our projects on our own, but that fallback strategy might make sense if somebody's really trying to.
Deep: Yeah, kind of get this, this,
Rob: there, there's, I mean there's a, there's a lot of different reasons and this makes our marketing hard, but there's a lot of different reasons people use the product, right? Sometimes it's cost improvements, sometimes latency improvements. Sometimes it's just like, what's the best model I can fail over to?
Rob: And then one of the trends we've seen recently is models are coming out faster, right? There's GLM five two now, and Kimmy models and all this kinda stuff is, people wanna know, should I consider that better model? How much better is it? So the way that we approach that is, we try to do as much as possible and we try to give the user our suggestions, plus let them easily add like the things they wanna try.
Rob: Because sometimes there's just, uh, people have a, these models are starting to brand themselves and sometimes people are like, nah, I wanna try the Claude models. And we're like, we don't think they'll do very well on this. They're like, I wanna try 'em. Or like, I'll give you an example. We recently tried, on the CR Marina data set just for [00:39:00] our own leaderboard.
Rob: We ran the, uh, IBM graphite models. I think they're called granite. Granite or graphite. I think it's granite actually. And, um, the granite models are a combination state space model plus transformer models. So they're a little bit different than what the Frontier Labs use. Uh, and they did not perform very well, but for certain use cases, they're very fast.
Rob: And so we are trying to expand more and more into these, unique models and unique opportunities and ways of doing things. And then, what we try to provide the user is just a list of options of like, Hey, you know, here's, here's some models that are faster. Here's the accuracy. Trade off for that speed.
Rob: You know, you'll lose 2% accuracy, but you'll get twice as fast. Do you wanna do that or not?
Deep: At the, at, at the end of the day. Like if you're, if you build trust with your callers and you get good at it. Isn't the end point that you just make the decisions for them?
Rob: We, we, we will get there. I mean, we could do that today, technology wise.
Rob: It's not hard, but I think to your point, it's, it's a trust issue, right? Which is, you wanna get to the point where, where people go, wow, you're making a lot of good recommendations. You know, we, we [00:40:00] should do this. And, there's this issue that you mentioned earlier in the, in the show about, people don't understand probabilistic reasoning.
Rob: So we also have to be like, you can't be too simplistic. Like, this is not a, like, tool for like, I don't know anything about ai, just help me pick a model. Like, we're not right for you. But also if you're super sophisticated, I mean, we've talked to a couple of public company teams that have built some super sophisticated infrastructure internally to like every task they have, they can spin up and test 50 models and do like, they're ahead of where we are and what we've built.
Rob: But, for those teams where I think a lot of the market is, which is like I'm a technical person. I mostly understand this stuff. I don't have the time to do everything myself, any of these tasks. That you can offload for me and automate report back. Let me understand. Like those are helpful so that's kind of what we target.
Deep: Yeah, I mean part of that just feels like the evolution stage that your company's in, where you're kind of getting traction right now and giving that feedback. But I think at some point, you know, you're gonna hit this point where people are like, just whatever, just optimize [00:41:00] it and just prove that whatever I was paying before it keeps going down.
Deep: You know?
Rob: Yeah.
Deep: Something like that. And then every once in a while they might come in and like, spot check you and test, but like, 'cause that's already happened with O Open ais already picking all of our models for us. Like we used to, you know, constantly I'd be picking models constantly and now I'm just like, screw it, whatever.
Deep: I don't care. You just pick. 'cause it's generally good. If I really need heavy reasoning all the time, I'll override. the one thing I wanted to kind of. Poke on a little bit is the issue of state. So, your kind of worldview makes a lot of sense when I can pass all the state information up with my API call.
Deep: But if I am, you know, like using a book or 10 books or like really long documents and I'm attaching those to, my API and now you have, now you're operating in a slightly more complicated space because each of the vendors have their own way of storing blobs, their own way of attaching them to the LLM session.
Deep: And now I'm sending in my prompts, but it's sort of divorced from however I configured my, [00:42:00] LLM session and. In text, it's one way, but like when you start getting into imagery and like, you know, it's another way. So how are you guys thinking about that and, and to what extent are you able to like, keep operating in this stateless way?
Rob: That is a really great question. So we are definitely better for now, like our ideal customer, mostly stateless has evals and everything else. But, but much like the eval question, we're, we're getting better and better at, learning to correlate, particularly for some of these age agentic work tasks where people have, multiple states and everything else.
Rob: We're working on some interesting stuff. Like, um, for example, if I know your system, can I make a, some kind of vector space embedding of the state of your system and pa like, let's see. 'cause you got a couple models through the, and I constantly pass that. Into a model or something like that to give a better output.
Rob: Like would you, if a model knew that it was in a system with a couple of other models and it knew the state of all the models, might it give you a slightly different [00:43:00] answer? So we're, we're working on some things to sort of solve this problem at scale because it's, it's not a huge issue right now. Most of the people that we talked to are relatively like, make a call, get a response, whatever.
Rob: But you can you can see us getting to a world relatively soon where you're like, oh, I've got four calls going out to four different models. And as they come back, I combine 'em and I do this to a frontier model. And then I, you know, it's, these things are gonna get more complicated. So we, we definitely have to answer this, but I, I don't know that it's super critical that we do it.
Deep: So, because, but the bulk of the calls you're seeing right now are still like the state, it's, it's reasonable for people to put all the state in the call.
Rob: Yeah. Yeah.
Deep: That's interesting because I'm not seeing that a lot in the stuff that we're doing of late. In the early days, it was like that where, where you were kind of doing like really narrow scoped reasoning, but a lot of the companies we're working with now, we'll have, like, they wanna really constrain the corpus or they really want to constrain the, the document that that, that they're extracting imagery from or whatever.
Deep: And so that means that you're trying to like, [00:44:00] take the entire LLM and like scope it down to this content, you know, either. So I was curious like if you're getting into their rag system and you know, plugging that in cleverly or anything like that.
Rob: We're not, we're not doing that yet, although so it depends, right?
Rob: There's a lot of ways to do this. Some people do like a lot of people using these agent builder tools. Part of the thing we do with those agent builder tools is in some of those tools, we are in the process of setting ourselves up as an endpoint that the agent builder can call. 'cause sometimes people want to use a tool like this, but they.
Rob: The models themselves are abstracted away from them. Right. And they don't, they know it's being used. Yeah. So, um, so that's, that's gonna be one way around that is working, with some of these infrastructure providers and different sort of partnerships. Uh, so
Deep: that makes sense. That's kind of like, that'll
Rob: help.
Deep: That's like a distribution channel for you guys.
Rob: Yeah.
Deep: Okay. So I want to, I wanna switch tacks a little bit. Um, so here on your AI injection, we like to kind of talk about three things. One is like what you do. I feel like we've covered that pretty well. Uh, the second thing is like kind of how you do what you do.
Deep: I think we've [00:45:00] probably covered that quite a bit. But the third thing we haven't really talked about at all, and it's really get, it's the should, should you do what you do. And I know it's a little bit hard to like wrap your head around in the context of your business 'cause it's a very narrow business, but let's assume for above that you're tremendously successful.
Deep: Mm-hmm. Do there exist any, like second order effects that you don't think about too much today? Or maybe you feel like you should. That could, have societal ramifications that are really bad. So like at a minimum, if you're successful, people are making a lot more API calls to LLMs and, and to, to AI system.
Rob: Yeah.
Deep: So I'd say at a minimum you fall underneath whatever the ethics of that, you know, winds up being. But like, are there any questions that you sort of ethically, like wonder about and think about?
Rob: Yes. Uh, I think about this a lot. I think we're a couple years away from knowing if this is gonna be a problem or not.
Rob: But here, here's an example of something I could think about that might end up poorly if we make, so right now everybody's scrambling and worried about like, oh my gosh, data centers and these things [00:46:00] use too much power. So that sets a certain high. R-O-I bar on using ai, right? Like, if it can't be this and solve this, then we can't do it If we make all these systems more efficient so that you're like, wow, you can use it for simpler stuff.
Rob: While your per usage price might drop, that might increase the demand already. 'cause it lowers the bar. For now you go, oh, well we wouldn't have made an AI model for this, but yeah, we should do it now. And that could have some really negative ramifications in the way that, I don't know if you ever saw the studies on Uber, but people thought like, oh, Uber's gonna make transportation so much more efficient.
Rob: It actually turns out that like in New York City where I live, right? You're like, um, before, well it's a little bit different 'cause there were a lot of taxis around. But taxis, it's hard to get a taxi if you're not on one of the main thoroughfares sometimes, right? Yeah. If you're in other parts of the city.
Rob: So what it turned out is like things that you would've walked four blocks and caught a taxi for, or walked seven blocks and then taken a subway. Now you just call an Uber right to where you're standing. So it actually has this weird induced demand [00:47:00] phenomenon where it's like, you take rides 'cause people did this static analysis where they're like, oh.
Rob: People take this many rides, and if we can replace those with Ubers and there's fewer cars and whatever, and you're like, oh, but what if Uber changes your calculation, second order effects on whether or not you take a ride. Well, if I could get the car to pick me up right here where I am, instead of walking four blocks and the weather's bad, it's cold or rainy or super hot.
Rob: So I think some of the analysis of Uber has been surprising and that it doesn't make traffic more efficient. Right. Because it causes people to,
Deep: well, it also has one person in the vehicle. Yeah. Right. Like, I think early on they kind of envisioned these things like taking over bus, like places the buses were inefficient and, and getting 3, 4, 5, 6, 8, 10 people in there.
Deep: Turns out people don't want to be, you know, in an Uber with eight other people they don't know. So like, or it's not as profitable for Uber. I don't know. So how does that relate to your business though? Like, what are you, what, what, what, what made you think of that Uber example?
Rob: So what you could see from our example is you could see.
Rob: People making [00:48:00] models for simpler and simpler things that they wouldn't have done before. And you could see that, you know, put putting, you know, people outta jobs who thought their jobs were more protected. You could see 'em using AI for more the way people may take frivolous rides. You could see people doing AI for more frivolous things 'cause it's so cheap that they might as well do it.
Rob: And those have negative impacts on society. Right. It's like, I mean you saw this with content back in Web 1.0. Right. Which was like, we were like, oh, there's all these gatekeepers. And they, they keep me from, they, they don't accept my letter to the editor, my op-ed piece or whatever. And I was like, well, we can all blog.
Rob: This is gonna be great. And to, to your point,
Deep: no,
Rob: we found some di we found some diamonds in the rough. We found some people who were great writers or whatever they
Deep: Yeah.
Rob: We also, but we also also had a lot of also
Deep: January 6th flood
Rob: and we
Deep: got
Rob: Yeah, exactly. Flood over the trash and, yeah.
Deep: Yeah. I mean it's, it's, it's anything, it's much harder to, to glean truth and how, you know, than it was 40 years ago.
Deep: Yeah. Yeah. So you, so, so your, your statement is basically like, Hey look, we are a cog in the efficiency wheel. We're gonna make this stuff more [00:49:00] efficient, more people, if we're successful, more people are gonna make more AI calls. Some of those AI calls could have been handled in other ways that are maybe much more efficient than using some LLM, which god knows, maybe it's an outer space.
Deep: I don't know. That sounds like bullshit to me, but like maybe the LLM is on earth, but, you know, using some emitting carbon, so let's jump up a level. Let's not talk about neuro metric. Let's talk about like, you know, you're in the space, you've been investing in AI companies for a while.
Deep: What's going on here? Like when we talk about, we can look at it from different lenses about like, what's the ethical ramification of this? Because when I look at the negative, like, AI is terrible, or AI is great narratives, they, they're all like radically oversimplified and almost always talking about the wrong things.
Deep: So I'm curious, like, what do you think are the right things that are actually real problems with AI in
Rob: Yeah, you make a great point because all the, I feel like mo 80% of the press is either like, oh my God, this is gonna solve every problem that we're gonna have a great life and we're [00:50:00] almost to a GI and it'll do everything for us and cure cancer.
Rob: Or it's like, it's gonna come alive and kill us all. And I, I think like you, I don't believe either of those, right? So I, I'll tell you a couple of theses that I have. Number one, maybe AI will hurt us someday. I don't, I don't know. I don't know that it has an incentive to do. I mean, I like to think about AI by comparing it to humans.
Rob: And there's two things that I think about. Number one, people are like, oh, well, the smartest human that, you know, think about who that is. Um, I can tell you for me, it's Steven Wolfram. I've met Steven several times. Oh,
Deep: I've met Steven too. Yeah.
Rob: Yeah. And
Deep: that I present to him, it was like one of the highlights of my career.
Deep: He is like, deep, that's good work. And I was like, oh my God.
Rob: Yeah, he's off the chart's. Brilliant. Um,
Deep: he's a break
Rob: cat. And, uh, I don't think Steven Wolfer could take over the world if he wanted to. Right. Like, I'm not even sure if somebody's smarter than Steven could. So, so the question was like, yeah, but these ads would be so smart.
Rob: Okay, but think about this. We don't know what it means to have a 400 or 500 or 600 iq. What we do know is every physical system has an asymptote, right. So it's like the [00:51:00] speed of light, the clo, the faster you go, the harder it is to go faster. Right. And there's a, there's a limit to the speed of light.
Rob: There's a limit to the speed of sound. There's. Limits to all kinds of things. Why isn't there a limit to intelligence? Why do we just believe it could just extrapolate infinite, infinitely everywhere? Um,
yeah,
Rob: so I, I don't believe there is, so,
Deep: I mean that the, the doomsayers around the smart box will take us over.
Deep: I, I find that just like, uh, frustratingly like shallow reasoning of being applied. It's like eventually, well, it's smart and it'll figure out how to jump outta the box. And it's like, yeah, okay. But, but there's so many real problems like staring at us right now. I mean, we have a president that had a bunch of people determine the tariff structure that made no sense to any economists on earth.
Deep: They had no idea what the hell this bullshit was that had come up, like when they came up with these tariff rates. And eventually somebody realized like, oh, some intern shoved this into GPT and this is what GPT spit out. And therefore it, like, it ends up like wreaking havoc on the markets. Now granted, like the, the chaos [00:52:00] agent was behind there, and would've exhibited in some other manner.
Deep: I mean, it just seems like there's so many like, real problems here and I'm just like, what do you think are the actual real problems? You know, like w with this stuff or like, what's the number one, one?
Rob: so because these ais can be connected to each other, as they start to run more parts of our lives, what are their decision making models?
Rob: So let me give you an example. Let's give a simple example. Let's use something like the AI algorithm that runs Waze, right?
Deep: Mm-hmm.
Rob: Which tells us where we go. Let's say Waze learns our patterns and it learns that every day you and I both go to work at nine o'clock, let get in our cars and drive, we drive this far whatever it has routes for us.
Rob: One day you're running early and I'm running late, right? Should Waze say, well, deeps early man, I'm gonna, I'm gonna make him delayed five minutes. 'cause he can bear it 'cause he's already running early so that Rob can get there on time. Is that the right thing 'cause our expectation is that Waze is making it as fast for us, so we spend as least time in the [00:53:00] car as possible.
Rob: Right. Or does it look at it and go, oh man, you know, deeps got a higher social rating. Or I, I can advertise this thing to deep that I can't advertise to Rob, that's worth more so I'm, I'm gonna make deep go faster. Or, or maybe I'm making spend more time in the car or like, whatever. Right. How, how do these algorithms assign value to the outcomes when you're dealing with lots and lots of people?
Deep: Yeah.
Rob: And you know, it's, it's
Deep: hard. I, I, I will generalize that a little bit. I call this the kind of the optimization transparency problem. So if you think about the classic kind of YouTube situation, or take any social media thing, you know, there a bunch of people who I don't think, wake up and think I'm Satan, I'm gonna destroy the world.
Deep: They're just like tech people who are like, whatever. They make a lot of money, they're doing their job. Somebody though, decided everything in this entire engine at YouTube is gonna be engagement optimization, right? So then the algorithms work, they're really good at figuring out how to optimize for one number.
Deep: And the engagement number is the one that [00:54:00] maximizes profit for YouTube, maximizes profit for the company. But what ends up happening is there's like a whole loss of transparency in all these black boxes down the line. Like all those nodes in the billion, trillion parameter network, whatever, black boxes.
Deep: Even in your system, if you guys are really successful, like all those, there's like a lot of transparency that needs to be elevated. Like, you know, we talked about how your customers, maybe they can define a ground truth. Well, maybe they can't. Maybe they just define it at the very end. Like, here's a 800 page.
Deep: Blob a text, give me this out. But they didn't segment it to all the internal steps that kind of produce that. And so when you have like some kind of number that you're optimizing for and then there's this, and there's an inherent loss of transparency because we as humans, we just can't really get in there and like figure out what's going on at the same rate or we just lose that ability to do so.
Deep: I think that that just leads to incredibly destructive outcomes [00:55:00] sometimes. Like YouTube being the great one where the, the algorithm figured out, this was like a few years ago, but like if you watched a Hillary Clinton speech, it just showed you something slightly weirder and then 20 clicks later you're like watching all these crazy left-wing conspiracy theories about, you know, nine 11 on the right.
Deep: Same thing if you watched a Trump speech, you 20 clicks down and you're like in. Complete like bullshit, crazy land. And the algorithm was great at getting you to watch all 20 of those videos. It did everything we asked of it, but we as society failed to like wrangle this thing. And I feel like that's our problem.
Deep: Like we are not wrangling this thing.
Rob: Yeah, I agree. You know, who has a good answer for this? Steven Wolfram actually has this idea that they should force you to, they should force social media platforms to create multiple algorithms and you choose the algorithm that you want, that gives you the results you want.
Rob: So you like pick from five and one might like, you should have an algorithm legally that's like, minimize my time on this app. Only show me the most important [00:56:00] stuff. Don't let me doom scroll. Right. And then it could be
Deep: chronological, it could be that's gonna be given to you in like all white font on a white background.
Deep: Like that option is not gonna be there. The same reason that that bloom and all these people making these cards to get you to stay off the social media crap. You know? I mean, God bless them. I use that stuff. I love it. But like, you know, it wouldn't surprise me at all if tomorrow they like sold to Facebook or something, you know, and then just got buried.
Deep: So, um, yeah, I mean that's like wolfman's idea's. Terrific. But like, can we at least just get the temporal, like the blatant timeline back like that we had 12, 14 years ago? Like, we can't even get that. Like, we don't, I don't even need a fancy thing, I just want the original timeline view and get rid of all this other bias, but no.
Deep: Yeah, it just, it's.
Rob: And then I think there's just a lot of problems with, like, I, I think most of all, I think it's overhyped how people do this. I think from, from the investor side, I think people, don't understand where differentiability and defensibility are gonna come from. And [00:57:00] I can tell you what I believe based on like watching the model labs and everything else.
Rob: I think there's this belief that like whoever gets to AGI first, even if it's just a matter of hours, they just, the thing self improves and it takes off and you can never catch it. I am willing to bet, and I'll say this publicly, it'll be the exact opposite. Whenever it happens, within three months, we will all have access to open source super intelligence.
Rob: And I don't know when that'll happen, but whenever we hit that threshold that people agree it's a thing, I think it'll be available to all of us. I, I don't know that anybody's gonna win and able to
Deep: keep else.
Rob: Well, that's
Deep: where Facebook, I think is being the good guys by like, open sourcing all the models as they come out.
Deep: But like,
Rob: yeah,
Deep: I don't know. I mean, do you think this whole a GI thing, like what is, is really like, what does it even mean as a, as a, as a terminology? Because I don't see LLMs taking us there. I see LLMs as sort of like hitting, like I feel like the efficacy, like the improvements are definitely slowing way down.
Rob: Yeah. Yeah. No, I, I totally agree. I think you're gonna need some different architectures to make that happen, and I don't yet [00:58:00] know what they are, but, but people are starting to get funded to do those. You're starting.
Deep: Yeah. I think Yann, Yann LeCun's view makes a lot of sense to me, this sort of objective based reasoning approach.
Deep: Like it could be years and I still think we have enough fodder here to like power 20 or 30 years worth of innovation, like just between the power of models we already have today. 'cause that, yeah, I, I think like if you look at the, the High Stack Foundation models and heavy tilt reasoning, it's pretty remarkable.
Deep: What that level of reasoning they're applying. And so, and I think they're like grossly underutilized still. Like we haven't figured out how to actually use all that. I think we're figuring it out really fast for software writing, you know? Yeah. Like, I mean, it's, it's really blowing my mind how improved Codex is, you know, over two weeks, three weeks ago or whatever before it came out.
Deep: So,
Rob: yeah.
Deep: But it'll, yeah, I don't know. I don't know what's gonna happen. So, but anyway, I think, um, thanks so much for, for [00:59:00] coming on the show. Is there anything that we didn't talk about that you kind of feel like we should cover, or?
Rob: No, I mean, this was awesome. Uh, we definitely could have talked a lot longer.
Rob: There's a lot of interesting topics here that we touched on, but, um, I, I, I think we hit the, I think we hit the crux of it.
Deep: I'll ask one last thing, 'cause I promised that I'd come back to your investment, um, background. So if you had to like lay down like one bet not on your company, but like in, and maybe not even in your space, but something else, like what's an arena that you think, it could, it could be ai, maybe it should be ai, but it could be something else, but like that, that you think is maybe not conventional wisdom, but something that really has a chance to be huge.
Rob: Ooh. Uh, good question. I would say one of two things. I think one of, part of what's being overlooked is, um, like industrial AI applications. So sensors on machines, un understanding how those things, you know, work together, where they're in process, how you can make factory floors more efficient and everything else, particularly as you get with robotics.
Rob: [01:00:00] And start to do more of that. And second thing I would say is, I still think. I think the hype in robotics is misplaced. I am less bullish on humanoid robots. I think there's still too many problems to solve. But I think on industrial robotics I think that is a space where with the current technologies we have in small, edge chips and the ability of some of these models, like you could take maybe a eight or 10 billion parameter model and run it on a small NPU, on the edge.
Rob: I think there's a lot of opportunity there. But humanoid robotics, you have battery power problem. You have the hand
Deep: problem. Yeah. I don't really get it. I don't, I mean, I just, I don't get why I think that's like, people have been reading too much sci-fi. Um, you can stick a hand on any kind of robot, right? Like just work on the hand.
Deep: Yeah. The hand's the part that's unique about the human. Yeah. It's not the fact that we walk upright, like Yeah. I mean it's the dog's running around like the dog thing's making a lot of progress from my, you know, chatting to my friends in the construction space. Like, I don't really get why, like, for example, why Tesla's putting all their, chips in the optimist bucket.
Deep: It seems weird to me. It just seems sci-fi worship, you know? [01:01:00]
Rob: Yeah,
Rob: I agree.
Deep: Maybe they learn something along the way and then eventually they, we, we roll out specialized bots that do whatever.
Rob: I, I, I think electric cars got too competitive and Tesla wasn't as dominant as they were and their way was to start to get outta that market and try to go somewhere else.
Rob: And Elon's good at telling stories. So that's my
Rob: theory.
Deep: Yeah. All right. Well thanks so much for coming on the show. I feel like, uh, this, I learned a lot and was really good and, yeah, and good luck with neuro metrics. I think it's, it's interesting stuff.
Rob: Yeah. Thanks for having me.

