Tackling Data Engineering Challenges with Autonomous AI Agents
In this episode, Frank La Vigne sits down with Pradnesh Patil, co-founder and CEO of Altima AI, to explore how AI is revolutionizing the world of data engineering. Together, they dive into the challenges of modern data stacks, the explosion of tools and technologies, and how AI-powered agents are transforming the way data teams build and maintain complex systems. From automating data pipeline management to optimizing infrastructure and safeguarding governance, Pradnesh Patil shares insights on the next wave of data engineering and what the future holds for professionals in the field. Whether you’re a veteran data engineer or a student curious about the evolving landscape, this episode offers practical advice, industry trends, and an optimistic look at how embracing AI can supercharge your career and your organization’s capabilities.
Links
- Pradnesh’s LinkedIn – https://www.linkedin.com/in/pradneshpatil/
- Watch on YouTube –https://youtu.be/vxQrjE-U3gw
Time Stamps
00:00 Evolution of data technology tools
04:11 Managing tool stack complexity
09:09 Discussion on AI guardrails and future
10:55 Managing MCB server outputs
14:32 Streamlining Data Pipeline Efficiency
18:48 Managing and Deleting Memories
21:12 AI security and virtual employees
27:16 AI transforming jobs and automation
30:43 Impact of AI on IT Industry
32:32 Upgrading legacy systems with AI
37:10 Importance of Data Engineering
39:10 Changing role of data managers
42:45 Home lab challenges and data management
Transcript
Something breaks and then you have to jump out of the bed in the middle
Speaker:of the night to get those things fixed.
Speaker:Hello and welcome back to Data Driven, the podcast where we explore the emerging
Speaker:industry that is AI, data science, and of course, none of
Speaker:it's all possible without data engineering. Now, unfortunately,
Speaker:my favorite, most favorite data engineer in the world can't make it here
Speaker:today. And I am actually enjoying the sunny but not too
Speaker:ridiculously hot sunny day here in the suburbs of Baltimore,
Speaker:Maryland. But I am excited here because I have Pranesh Patel,
Speaker:who is the co-founder and CEO of a company that
Speaker:does— makes data engineering a lot more palatable, it sounds like, Ultima
Speaker:AI. Welcome to the show. Hey, um,
Speaker:Frank, thanks for having me here. Super excited to chat with you
Speaker:today. Yeah, so, so tell me about your company, Ultimate AI. Is, uh,
Speaker:tell me about what it is and what led you to make it.
Speaker:So Ultimate AI, we are a startup based out of the San Francisco Bay Area.
Speaker:What we do is we use AI to
Speaker:automate a bunch of data engineering tasks for the teams out there
Speaker:who do data work, right? The tasks like building
Speaker:ELT pipeline, extract and loading of the data, transforming
Speaker:data, and not just building those pipelines but maintaining
Speaker:those as well. It's always nightmares. Something breaks and then you have to
Speaker:jump out of the bed in the middle of the night to get those things
Speaker:fixed. So building and maintaining data pipelines or
Speaker:managing your data infrastructure Like in the modern data
Speaker:stack today, we have Snowflake, Databricks, BigQuery, those things, right?
Speaker:We need to spend a bunch of effort to manage that infrastructure as well as
Speaker:optimize that infrastructure, like making sure you have turned on the right knobs on the
Speaker:infrastructure side as well as you have optimized those queries and pipelines,
Speaker:etc. So we work on that as well. The whole idea is how we
Speaker:can use the agents to automate a bunch of that work. so
Speaker:we can scale our teams even further. I'm glad
Speaker:you mentioned— I'm sorry, go ahead. No, you asked
Speaker:me that, hey, how did it all start? So me and my co-founder,
Speaker:we have worked in B2B enterprise space for a long time. You must have seen
Speaker:all those data teams every company has. There's so much work involved. There is a
Speaker:long backlog. We went through the same experiences, and we started
Speaker:thinking there must be a better solution so that we can handle that
Speaker:backlog. At the same time, a few years ago, this whole AI wave started
Speaker:coming in. We jumped in headfirst and started building
Speaker:agents first to sort of automate simple tasks like writing
Speaker:documentation or writing some data quality tests. And through that,
Speaker:the tool and the product evolved from there. And now
Speaker:we automate bunch of things that are day-to-day things for
Speaker:data engineering folks. That's a good point. You brought up a good point
Speaker:because you mentioned Snowflake, you mentioned Databricks, you mentioned all of these
Speaker:technology tools that the stack used to be a lot simpler in data,
Speaker:in the data space, right? You, you know, you were either an Oracle shop
Speaker:or a SQL Server shop and you had, you know, SQL
Speaker:Server tooling, which those I'm way more familiar with, but also the
Speaker:Oracle kind of stack too, right? So it went from being kind of a
Speaker:The tools may have been limited, yes, but the tools also—
Speaker:there only wasn't really that much of them, right? You picked one side
Speaker:and stuck with it, right, for most organizations. And then we
Speaker:had things like Hadoop and Pig and Hive and all of those things kind of
Speaker:come out. And now this many decade and a half or so
Speaker:or more, you can tell by my gray hair, that now we
Speaker:have just an unlimited assortment, it seems, of these
Speaker:tools, right? And each one of them has their own quirks, their own settings, as
Speaker:you said, the knobs and dials. Is that something that your
Speaker:solution offers a solution to? Yeah,
Speaker:absolutely. Because as you mentioned, the use
Speaker:cases exploded, and with that, the tool stack exploded as
Speaker:well. You can't be exploiting 35 different tools to make sure
Speaker:everything works perfectly and keep your tabs on everything.
Speaker:So we help tremendously with that, like all those configurations in your
Speaker:infrastructure as well as building out things. You can't be expert in so many tools.
Speaker:And on the other side with AI, what has happened is
Speaker:even these tools as a software are evolving so fast. In a
Speaker:year, there are like 100 features come out. How can you keep tab on what's
Speaker:the latest and greatest and what actually fits in my environment and to my use
Speaker:cases? And, you know, AI for Rescue there, it can
Speaker:do a bunch of those things for you. No, absolutely. And
Speaker:I hadn't really thought about that angle of you have to be an expert.
Speaker:It's like what happened in the software development world
Speaker:where you had the notion of a full-stack developer where suddenly,
Speaker:you know, it went from you would be a Visual Basic developer, right?
Speaker:Or a developer of, you know, web developer. But
Speaker:then it became, no, you had to be a full-stack developer. You had to know
Speaker:the data side. You had to know the CSS. You had to know the HTML
Speaker:and the JavaScript, right? Yeah. Data, I think, has followed a very similar trajectory
Speaker:in that regard of no one person can do it all. At least
Speaker:no one person can do it all without the assistance of some kind of
Speaker:AI. And is that what your product enables? Like
Speaker:somebody to like basically be a— in the DC area,
Speaker:they love the term force multiplier. Is that kind of like what
Speaker:your tool does? Exactly. Exactly. Because
Speaker:as I think things exploded, use cases exploded, and even
Speaker:the use of AI has exploded overall, your data
Speaker:actually fuels all of it. And what's happening is the data projects
Speaker:are growing exponentially, but our teams are not growing in that size. And
Speaker:that's why there is a huge backlog that's happening. So the idea is then,
Speaker:in your words, how we can have some sort of force multiplier that
Speaker:can do a bunch of things for me automatically. And I focus on some of
Speaker:the more complex things or something where AI needs more help
Speaker:and guidance. So think of this as these bunch of agents
Speaker:are my minions. They're getting the work done. I have like 30 of those and
Speaker:I'm giving them directions, course correcting, telling them what to do. And
Speaker:that's how basically the teams are scaling themselves today. But there are so many challenges
Speaker:in doing that as well, in that whole process. And our
Speaker:goal is how we can streamline and make that whole process smooth. So you have
Speaker:teams of these minions which are getting a lot of work done for you.
Speaker:Yeah, and I think you brought up a very real pain
Speaker:point, right? They're not hiring tens of people in the data engineering
Speaker:teams anymore, right? Data engineers are expected to do, to all
Speaker:be 10x engineers, right? And so
Speaker:I like the idea of you calling these AI agents minions, one, 'cause
Speaker:I think the movies are really cute and I have small kids. But
Speaker:how do you— what's the governance look like on that? Is that something that you—
Speaker:how did you address that problem? 'Cause I'm sure that comes up quite a bit.
Speaker:No, I think you touched on very important area, right? Because as a
Speaker:human, we usually have general sense of understanding and
Speaker:what needs to be done and what shouldn't be done. But for agents,
Speaker:it's a piece of code, it's a machine, right? A lot of times it
Speaker:doesn't have that, compass of common sense. So for example,
Speaker:that's why you might have seen situations where agent ran a query and that cost
Speaker:you thousands of dollars. Oh, and there's so many stories of agent
Speaker:deleted my sensitive data or replicated it. So whether it's sensitive
Speaker:data, access controls, cost guardrails,
Speaker:all of those are very important factors around which
Speaker:we need to develop a solution. Now we just talked about Lot
Speaker:of tools in the data stack. Every tool has their access layer, RBAC
Speaker:layer. Now, how are we going to do this, right? So what we have been
Speaker:doing is, and that's the reason we launch the, our product as open
Speaker:source project, we call it Ultimate Core. There is a governance
Speaker:layer that's built in, which allows people to define these guardrails, rules,
Speaker:and permissions. Lot of guardrails come inbuilt.
Speaker:And the beauty of that is this is a layer that sits on top of
Speaker:these bunch of different tools that you are using in your data stack. and gives
Speaker:us the common ground around governance. So for example,
Speaker:cost, it makes sure the agent doesn't spend more money than this
Speaker:for your specific task. Or for agent
Speaker:itself, you can create a separate view, separate tables so that they don't
Speaker:inherit like service user permissions or user permissions directly. Because as
Speaker:a user, I might have so many permissions, but I don't want my agent to
Speaker:have the same permissions and use my same credentials, et cetera.
Speaker:So that part also we have solved pretty well.
Speaker:Yeah, I think, I think calling them minions works out pretty well too, because, you
Speaker:know, the minions always— there's a whole sequence in the first movie about the dart
Speaker:gun. Parents, if you know, you know. But
Speaker:the minions misheard it and built something else. But
Speaker:no, I think you're right. Making sure these things have guardrails on them,
Speaker:I think, is— people are going to learn very quickly what happens if you don't
Speaker:have guardrails, right? millions of dollars a query, or etc., etc.
Speaker:Actually, recently in the news, and they haven't disclosed what was the core— what
Speaker:was the core problem, but apparently AWS had—
Speaker:was giving out bills that were like orders of magnitude higher.
Speaker:And I can only— again, I have no inside information,
Speaker:but I can only imagine that there was probably some kind of rogue AI doing
Speaker:a math error or doing something crazy like that. What do you think the
Speaker:future of these— this tool space is going to be? Do you think You know,
Speaker:there'll be more MCP servers, right? What do you think,
Speaker:um, what do you think that's going to look like, the ecosystem,
Speaker:the data ecosystem? So how I
Speaker:see it is based on what we want to do, the tools
Speaker:get built, right? And as we sort of expand our
Speaker:horizons to do more and more things, the tools get expanded also, right?
Speaker:For example, I think MCP server is a great example. At one point
Speaker:we realized, hey, we need to sort of feed in all this
Speaker:information to MCP server, uh, to AI agents, right? And then
Speaker:MCP servers were built as a solution to it. Now
Speaker:we are hitting the limits of MCP servers themselves where,
Speaker:you know, okay, governance— anybody can install any MCP server and I have
Speaker:no control over who uses which MCP server in the organization. Right. That's
Speaker:happening. Second is The MCB servers are putting out so
Speaker:much output that my tokens are getting consumed like candy,
Speaker:right? So then how do I control it where MCB tool outputs are
Speaker:limited? And we have built some functionality around that, around context comparison
Speaker:and tool output curtails, etc., especially for data engineering
Speaker:tasks. Our third big thing is MCB server is like
Speaker:API, is going to just pull the data from the other system and bring it
Speaker:to you. But it is not going to do any intelligent filtering
Speaker:or connecting those dots together. And then now people are looking at
Speaker:semantic layers and those things, how I can do it also. So now we are
Speaker:running into the limitations of NCP servers themselves, and then people are trying
Speaker:to figure out, hey, what's the next thing we need to do to solve this
Speaker:problem really well? And then people are talking about context graphs
Speaker:or context store where all of your information is already there,
Speaker:curated, filtered, smartly arranged, And then you use MCP
Speaker:server to pull only right information instead of directly interfacing with the
Speaker:tool, like for example Salesforce or HubSpot, and just dumping everything into your
Speaker:AI agent as well. So how I see it is, I think
Speaker:agents are going to become more and more autonomous, more and more ambient
Speaker:as well, and the information we are going to feed them
Speaker:is going to become much more curated as well. And there are so
Speaker:many facets to this. There is a right information feeding angle around
Speaker:context. There is a governance angle to it, and now. And now I think you
Speaker:just touched on it. The one big angle that's coming into the play is cost
Speaker:as well. There, there have been so many stories coming out where people spend
Speaker:their entire year's worth of budget in like 3 months. I know some stories
Speaker:which are not public where people's usage like 30x'd in
Speaker:6 months and now they're like, oh my God, my AI bill is actually same
Speaker:as my cloud bill now. I never planned for this.
Speaker:And what is exactly the ROI that people are trying to measure also?
Speaker:Yeah. So all these questions are coming up, and as the questions and use cases
Speaker:come up, I believe we'll have more tooling and better solutions.
Speaker:Is anything— is there anything in particular that your product addresses
Speaker:to any of these problems? Yeah, so we talked about the governance
Speaker:piece of it. On the cost side as well, what we started doing
Speaker:is one of the features we have in the product is context compaction.
Speaker:So I was talking about MCP tools. Dumping a lot of data, like
Speaker:especially data-related MCP tools. So what we do is
Speaker:we curtail the output correctly. We know all these MCP
Speaker:servers, etc., so that you are not spending too much of a token cost. And
Speaker:not just the cost, right? What they call is a context rot. If you
Speaker:dump in too much information, LLM has too many directions to go in. So we
Speaker:curtail and do that context compaction. Second part we
Speaker:have introduced is Specifically for data tasks, we create
Speaker:memories automatically. So not every time agent is starting from scratch,
Speaker:and it's very curated for data-related tasks. So in that way,
Speaker:next time when the agent does the same task, it can tap into those memories
Speaker:that are shared across the organization and get the task done in
Speaker:very less number of steps. That's an interesting
Speaker:point because I noticed that, that was one of the When I started
Speaker:poking around my OpenClaw instance, right, inside of there, there's the
Speaker:soul, but there's also kind of this memory type of notion and managing
Speaker:that memory so the context doesn't have to start from
Speaker:zero every time. And does that really save on tokens? Like, what's
Speaker:the rough order of magnitude in terms of what the token save is?
Speaker:No, the memory will save you tremendously because, for example,
Speaker:let's take an example of data pipelines. Right? So if your data
Speaker:pipelines are failing, usually there is a pattern, same issues you're going
Speaker:to see again and again. Hey, data hasn't landed, that's why this pipeline particularly
Speaker:fails. Now if that thing gets stored in the memory, your agent
Speaker:is going to check the first thing is has data already
Speaker:landed? That will save you other 10 other things
Speaker:that agent would normally try out before coming to that, right? Boom, your
Speaker:multiple workflows are saved. Your token cost saved, your time is saved
Speaker:also. We fixed it very quickly, right? So it helps us
Speaker:tremendously in that way. And the beauty of that is even the memory
Speaker:layer cannot be very generic, right? You need to understand
Speaker:as tasks are happening, what kind of memories are important. It needs
Speaker:to be curated. There needs to be, say for example, somebody's writing
Speaker:agents and harness for finance, there needs to be a finance-specific memory. that
Speaker:will remember finance-related things. Similarly, for data work, what we
Speaker:have created is data agent-specific memories, which works amazingly
Speaker:well. And since we are talking about memory, right, memory
Speaker:is not just limited to one session. So for example, what I'm trying to say
Speaker:is, now if you have data pipelines, probably you have a team of
Speaker:engineers maintaining that data pipeline. So for example,
Speaker:I fixed this data pipeline today, for that data not landing
Speaker:issue, and it gets saved in my memory. Of course, next time
Speaker:in the regular systems, I try to fix it. Next time it will read from
Speaker:the memory stored in my, say, code editor or Cloud Code or something like that,
Speaker:and I can draw to it. But what about somebody else on my team? They
Speaker:are also going to work on that pipeline, and they might be on on-call. That's
Speaker:when the pipeline failed. So the memory layer that we have
Speaker:built, it's a layer that gets shared across the teams and
Speaker:organization as well. So then it's the— I call this a
Speaker:hive-like mind in which you are storing the memory, and anybody can
Speaker:come in and use that hive-like mind and sort of use
Speaker:that knowledge to do things better. And the big benefit of this is
Speaker:even for newer people joining your team, right, they don't have to start from scratch.
Speaker:All this tribal knowledge is stored in that hive mind, which
Speaker:we call as a memory layer. I like that because then
Speaker:your first few rounds of learning it, right, you're
Speaker:literally onboarding this virtual employee, it sounds like, right? It sounds somewhere between a minion
Speaker:and a virtual employee that can kind of capture that tribal knowledge,
Speaker:capture that institutional kind of wisdom. Yeah. And so
Speaker:it's not so much you're spending tokens, you're kind of investing tokens, right,
Speaker:for the future and training this virtual employee. And
Speaker:presumably you'll get that, you'll see dividends later on.
Speaker:Exactly. And this system works with agentic
Speaker:frameworks that are out there already. People use, say, Claude Code, GitHub
Speaker:Copilot, Cursor. The system works with those
Speaker:already. We don't build LLMs, we don't build—
Speaker:give agentic frameworks, but we build this harness which makes
Speaker:these tools extremely suitable or powerful when it
Speaker:comes to data engineering work.
Speaker:Interesting. Can— does it learn on its own or
Speaker:can you edit those memories? Right. So like, what if— what,
Speaker:let's just say we have a, somebody makes a mistake, right? You obviously
Speaker:wanna mark that and kind of remove that from the memory. Is that, is that
Speaker:like, obviously you probably edit, it probably adds its own. And,
Speaker:and the reason why I mentioned this is because when I dove into my
Speaker:my OpenCLAWS memory file. I thought it was funny what it read about me.
Speaker:So if people are watching this, you'll see I'm kind of like winking my eyes.
Speaker:Apparently allergies are really bad today, which I did not factor that in when
Speaker:sitting outside. And one of the things it learned about me was I
Speaker:always ask about the pollen levels for the day, which I think is kind of
Speaker:funny. It said that, you know, I ask about the weather, I ask about stocks,
Speaker:and I ask about, you know, AI innovations and pollen report,
Speaker:right? Does this— but obviously I can go in, I can open up a terminal
Speaker:and edit it myself. But how does your solution— is it built into
Speaker:the UI or is it like just a config file? So
Speaker:as you start using the solution, you install, it will start creating
Speaker:memories automatically. You can, of course, in your prompt give a
Speaker:specific instruction. As you give the instructions, those get saved also. But you
Speaker:can say specifically create a memory also. And through our
Speaker:MCP, it will get that memory created as well. Now the harder
Speaker:part usually is what if there are bad memories and you want to change memories
Speaker:or you want to erase those out? The good news is it comes with that
Speaker:flash stick that I think I remember it from the movie Men in
Speaker:Black, right? It comes with that. So you can go in, delete your
Speaker:memories because as I was telling you, we store these memories in a SaaS. So
Speaker:in that way, they're shared with your team members, they're shared with the rest
Speaker:of the organization, and there is a granular control you can do who they get
Speaker:shared with. But at the same time, if you want to update
Speaker:it, you can go to the UI and get those updated or
Speaker:deleted as well. Oh, interesting. So you
Speaker:have that neuralyzer built in. I think that's what the thing is called in Men
Speaker:in Black. But so you can go back
Speaker:and you can remove like, hey, everything I did today was terrible, so don't remember
Speaker:that. What about
Speaker:security, right? Obviously, There's a lot of trade, you know,
Speaker:there's a lot of sensitive information that are going to be floated around in the
Speaker:data engineering space. How does your solution, how does Ultimate AI
Speaker:kind of address that? So first and foremost,
Speaker:we don't look at the data directly. Okay, so it's a
Speaker:harness that comes in, and this harness people can install locally
Speaker:as well. So if you're some sensitive industry, maybe healthcare
Speaker:or something like that, You can use it completely
Speaker:locally, and the LLM solution that you use in the
Speaker:background, people can use their LLM solution. As I was talking about, Claude
Speaker:Core subscription or Codex subscription, they can use
Speaker:that. They can hook up their own models also. We support, say for example,
Speaker:OpenRouter, all these different models people can use. Even if they want to use
Speaker:on-premise LLM, we support that as well. On the other side,
Speaker:as a company and as a platform, We are SOC 2 certified,
Speaker:pen tested, a bunch of big Fortune 500 companies
Speaker:use our product already. So even on that side, I'm
Speaker:sure we can make people's security teams happy because we have done a bunch of
Speaker:work around it already. Oh, that's interesting. That's good.
Speaker:Because I mean, the security conversation
Speaker:comes up in AI, but I don't think it comes up often enough or early
Speaker:enough. And Obviously, I think that's going to change as more and
Speaker:more systems get deployed. And obviously, the Fortune
Speaker:500 companies obviously also take that into account as
Speaker:well. What would be your advice to people who are data
Speaker:engineers who are curious about how do I make— how do I get my own
Speaker:minions, right? Like, how do I start thinking about— let's roll
Speaker:that up. How do I start thinking about
Speaker:my job as a data engineer in terms of
Speaker:I'm managing a dozen potential virtual employees
Speaker:as opposed to doing it myself, quote unquote, the old-fashioned way.
Speaker:Yeah, yeah. No, I think that's how we should start
Speaker:thinking about it if anybody hasn't started going on that
Speaker:path. There are a bunch of tools out there, and I
Speaker:believe it's easy to get lost also because there are just so many tools and
Speaker:things that are happening. on the AI side of the things. And what I have
Speaker:seen is people try to retrofit the tools for software engineers to data engineering,
Speaker:and usually that leaves bad taste in their mouth. And sometimes
Speaker:I heard, hey, AI doesn't work. It's because you're using the wrong tools for data
Speaker:engineering most of the times. So you need to use the
Speaker:harnesses or tools that are specifically done for data engineering.
Speaker:On our side, we open source Ultimate Code Project, so definitely try it out.
Speaker:It's MIT licensed, completely open source. But in addition to that, what
Speaker:we have done is on our website, we have put this
Speaker:course called AI Data Engineer. So there is a section on our
Speaker:ultimate.ai website, AI Data Engineer, where we have put
Speaker:curated articles and we keep spending time to make sure all the updated
Speaker:information is there. So somebody might be like at day 1, somebody might
Speaker:already be at day 50. You go through that material and it will coach
Speaker:you. What are the latest tools and what are the best things you can possibly
Speaker:use to get these things done? Oh, that's really cool. So you
Speaker:have built-in training, and you did mention that your product is open source. We'll make
Speaker:sure that the link to the GitHub repo and as well as the, um,
Speaker:this onboarding, um, this kind of this ment— what would you call it,
Speaker:a mentoring tool, an onboarding tool? Like, what would you call
Speaker:that? I just call it training boot camp. So
Speaker:join it. Even if you're brand new, there is a bunch of useful
Speaker:information. Even if you have been using it for a while, you know, world
Speaker:of AI and technology changes so fast. Definitely it will help you
Speaker:keep tabs on what are the things changing as well.
Speaker:Yeah, no, I— it's a, it's a very fast-moving space
Speaker:and it, it's not gotten slower. It's— I think the pace of innovation is
Speaker:accelerating for good or for bad. What
Speaker:What would be your advice to someone? I started asking this in the question, like,
Speaker:there's a lot of computer science students that are very still
Speaker:in school and they're very worried about AI taking their jobs. I think you and
Speaker:I have gotten to the point where AI doesn't really take away your job. I
Speaker:think it changes the nature of your job.
Speaker:Yeah. But what would be first? My advice would be stick it
Speaker:through, kids. But like aside from that, what would be
Speaker:your advice to particularly, so I think a lot of people
Speaker:are getting started in data engineering and if they hear that there's AI tools
Speaker:assisting with that or doing that, they may be a little bit afraid or concerned
Speaker:about the future here. I think the future looks bright, but that's just my opinion
Speaker:as a, I wouldn't say I'm a perpetual
Speaker:optimist, but I've kind of been through that cycle of,
Speaker:oh, you know, I, AI is taking away my job, but I really kind of
Speaker:see the nature of it. It just changes the nature of my job.
Speaker:Yeah, yeah. If I can summarize this, I have heard those
Speaker:things, hey, is AI going to take my job? And I always tell
Speaker:if anybody says this to me, in my opinion,
Speaker:it's not AI that's going to take your job, but it's going to be somebody
Speaker:who can use AI is going to take your job. So if you don't upscale
Speaker:yourself, because I completely resonate with your point. Our
Speaker:jobs and roles are changing rapidly in this new world, so
Speaker:upskilling yourself for AI is very, very important. And being
Speaker:on that cutting edge of the AI and data engineering and all that data
Speaker:work, that's the most important part. How I'm feeling
Speaker:is going to unfold is like how 100, almost 100 years
Speaker:ago, Industrial Revolution unfolded, right? There were, for example,
Speaker:there were people who are making clothes by hand, right?
Speaker:Then the machines came and we got operators who would operate those machines
Speaker:to produce clothes at a humongous scale and see how
Speaker:many clothes we have today, right? Same thing is going to happen
Speaker:with AI in general. It's going to give us those machines
Speaker:which will increase the overall throughput of all the people
Speaker:out there for different professions. It's going to bring in a lot of prosperity.
Speaker:But at that time, people who were making clothes by hand, they went out of
Speaker:job. Maybe some of them learned how to operate the machines, but that's
Speaker:the important part. Learn to operate the machines. Like, don't just say that, hey,
Speaker:AI is going to take my jobs. Learn to use AI, embrace it. Your job
Speaker:is not going anywhere there. That's a really good way to put it.
Speaker:One of the examples I like is, um, there's self-driving cars,
Speaker:right? But there's also a lot of cars have adaptive cruise
Speaker:control. So for those not familiar with it, you know, it's basically cruise
Speaker:control, which will keep your speed, but it also has proximity sensors. So
Speaker:it'll apply the brake, right? And kind of keep you within a certain distance of
Speaker:the car ahead of you, et cetera, et cetera, et cetera. And also can keep
Speaker:you inside your lane. When I got that in, in
Speaker:my car for the first time, I felt like— It's a good analogy for AI
Speaker:because I'm still controlling the car. Right. I'm not fancy enough
Speaker:to have a full-on, you know, self-driving car, but, um, but I
Speaker:did notice that I became more like a captain of a ship, so to speak,
Speaker:where I would— I felt like I was guiding the car which direction to
Speaker:go. Now, in that case, I'm not a big fan of driving around
Speaker:all the time, but so it would be nice to have a full driving car.
Speaker:But I think that's kind of a good analogy for how,
Speaker:how jobs are going to look like, right? And if you're a software engineer, you're
Speaker:not going to be writing every line of code by hand anymore. I think,
Speaker:to your example, I think, you know, when humans were sewing, doing the
Speaker:sewing, right, individual stitch by stitch, clothes were made. And accordingly,
Speaker:clothes were expensive. And I think we're going to see that kind of happen with
Speaker:AI, right? Whether it's code, whether it's data engineering tasks.
Speaker:Because certainly I think anyone out there who is in a professional
Speaker:position where they're doing data engineering, There's plenty of
Speaker:things that are just kind of on the back burner that they would do to
Speaker:make their jobs more efficient, right? And I just think that,
Speaker:I think by having AI enabling,
Speaker:you have a lot more cognitive free room. And all those things
Speaker:are on the back burner, on the whiteboard somewhere, or kind of on a Post-it
Speaker:note somewhere about, hey, you know, I bet I can improve this process by
Speaker:X number of percent, but they're too busy keeping the machines
Speaker:working, right, keeping everything working. So with AI, I think you can have a
Speaker:bit of cognitive surplus where you can tackle those. Sorry, I cut you off. You
Speaker:were about to say something. No, nothing to add there. I think
Speaker:that's so spot on. This is where the world is going.
Speaker:That's cool. That's cool. So what's next for your company? Like, what,
Speaker:what's, like, what other big challenges do you think you
Speaker:all are going to take on? So one thing is, of course,
Speaker:right now we have been automating few of the data engineering use cases,
Speaker:but we are planning to expand to more and more use cases and at
Speaker:the same time improve the support of the data stacks
Speaker:we support as well. For example, you're talking about SQL Server,
Speaker:maybe older Hadoop, Hive stacks. There's so much information
Speaker:out there, and being early-stage company, we support certain
Speaker:stacks, certain stacks we don't support. So the plan is to expand
Speaker:into more ecosystems and at the same time expand for more
Speaker:use cases as well. Yeah, that's cool. Yeah, and it's funny
Speaker:because I know there's a lot of Hadoop legacy solutions out
Speaker:there. Somebody told me, and I don't want to pick on Hadoop, right? But
Speaker:Hadoop was one of the first, you know, big data solutions out there. And
Speaker:I would imagine there's a lot of— somebody told me that it was moved to
Speaker:the— what's Apache call it? The attic? the attic or the basement where it's basically
Speaker:considered a legacy product now. And, you know, obviously
Speaker:there's a lot of engineering— data engineering projects historically
Speaker:have not turned over that quickly, right? Like, you know, there were— there are probably
Speaker:tch jobs I wrote in the early:Speaker:And that's not that unusual. So I think that
Speaker:because data modernization projects even if they're
Speaker:needed, they may not happen because people are too busy keeping
Speaker:the machines running, right? So if you can kind of— I think people are missing
Speaker:the point about AI taking their jobs. I know we're back on this again. You
Speaker:can always tell when the coffee hits me, right? Like, my mind goes in like
Speaker:3 different places at once. But the release
Speaker:of cognitive service, cognitive surplus,
Speaker:I think is going to do a lot of good things for The IT
Speaker:industry, and ultimately, I think, for the broader economy as a whole, right?
Speaker:I think you, you mentioned prosperity before, and I know it's a, a lot of
Speaker:people are very anxious about what AI can do. And you live, you know,
Speaker:you're in the Bay Area, you're probably at the epicenter of a lot of this
Speaker:back and forth, right? But I think that
Speaker:by releasing some of this, the human minds
Speaker:from the drudgery of some of the, the machinery that we have today,
Speaker:I think has enormous potential for improving IT, right?
Speaker:You know, take any kind of like regular person out there who, who has to
Speaker:interact with their IT department. Would they describe it as a wonderful experience, right?
Speaker:Would they describe it as a, you know, or is it more like going to
Speaker:the DMV, right? Probably more like the DMV. Although
Speaker:I will say credit where credit is due. In Maryland, I had to spend— I
Speaker:recently got a new car and it wasn't as bad as I feared
Speaker:it to be. You know, maybe they're using AI, I don't
Speaker:know. Yeah, yeah. And as you mentioned, right, a lot
Speaker:of this work we sort of parked it for later. Like you were talking
Speaker:about batch jobs or Hadoop jobs, etc. There's still a bunch of those things running
Speaker:because migrations had been nightmares, right?
Speaker:Like people spend millions for years. But now I think, now
Speaker:if I look at it with the angle of AI, AI can automate
Speaker:90-95% of those migrations as well. Now, who knows, like, even
Speaker:that whole modernization and digitization
Speaker:in the data world, that will start accelerating much faster because of
Speaker:this. And a lot faster, different companies can be on the
Speaker:most modern technology. And they can, they can take advantage of
Speaker:the newer features, the optimization technology. Yeah, I mean,
Speaker:because no, no company in their right mind is going to suddenly, you know, say
Speaker:we're going to hire, you know, we have this system, we can either
Speaker:A, keep throwing a little bit of money at
Speaker:it per year, right? And keep it going indefinitely. Or
Speaker:B, we're going to hire 100 new data engineers and all these project
Speaker:managers and things like that. We're going to spend billions on upgrading it. They're not
Speaker:going to do that, right? But they do need to upgrade, right? So then if
Speaker:it becomes a, well, you know, if you just spend just a little bit
Speaker:more than what you do have to do to keep the lights on, and you
Speaker:kind of have AI take a lot of the brunt work of that, right? Instead
Speaker:of hiring 400 people, you can probably hire, you know, say
Speaker:40, right, to do the migration. It starts to become more
Speaker:palatable, right? It starts to become more justifiable in a, in
Speaker:a balance sheet. And, and, you know, I think a lot of the breaches
Speaker:we've seen have been from improperly
Speaker:configured systems, right? Or legacy software
Speaker:that hasn't been patched or legacy software that has no business still running
Speaker:in this day and age. And I think you're right. I think AI is really—
Speaker:I think if you put the right type of glasses on, AI
Speaker:is a way to solve the problems that we've just
Speaker:been pushing back for years and years. And technical
Speaker:debt, that's the word I'm looking for. I think AI has the potential to
Speaker:be a great way to pay down technical debt. Yeah.
Speaker:But So while we're talking about
Speaker:that, like, what is the thing that your customers who are successful on your
Speaker:platform, what is the thing that they're most happy with at the end
Speaker:of the day? What are they like, they call you up and they say, wow,
Speaker:Pranesh, I'm so glad we did this because
Speaker:what is it, increased productivity? Is it something else?
Speaker:A few things, right? So for example, one example I gave was
Speaker:we can manage people's infrastructure using AI. Right? And
Speaker:we can optimize it. Now AI does it at scale
Speaker:at which humans can't even do it. For example, it can change the configurations
Speaker:300 times in a day. We can't do it, right? Like, how can we
Speaker:change the configuration of single machine 300 times a day? You need 150
Speaker:people doing just that. Yeah. And
Speaker:especially in the space of data, your workloads are always changing, right?
Speaker:As your data changes, your workload changes. So one of the things we have done
Speaker:is that infrastructure optimization, where we have built agents which can
Speaker:automatically analyze your infrastructure, change the configurations. We have built agents
Speaker:which can analyze the data pipelines and automatically optimize them. And
Speaker:of course, that saves tons of time for engineers, but at the same
Speaker:time, your infrastructure runs at 90 to 100%
Speaker:utilization, and those pipelines are always top-notch when it comes to
Speaker:optimization because people run millions of SQL queries and hundreds of thousands of
Speaker:pipelines. Now AI does that. The result is not just the
Speaker:engineering hours saving, but real dollar savings on
Speaker:the infrastructure cost also. So some of our biggest customers, we
Speaker:have saved them millions of dollars in their Snowflake and
Speaker:Databricks environment by optimizing this.
Speaker:And yeah, that's super happy about it. They're like, hey, tool pays for
Speaker:itself multiple times over. We are getting all these engineering productivity also. We're saving
Speaker:tons of engineering hours. But hey, right there and then you're saving me
Speaker:infrastructure costs also by automating that process by a huge margin.
Speaker:No, it's a great way to— that's a great way to look at it. And
Speaker:I really think there's so much opportunity. I know there's a lot of people, there
Speaker:are a lot of naysayers now about what AI looks like and,
Speaker:and, you know, will we, will we realize the gains that were been promised?
Speaker:I think we will. I think it's going to be in places where we may
Speaker:not think of, right? No one You know, and data engineers will be the first
Speaker:people to tell you, like, they're not usually first of mind for a lot of
Speaker:people in the C-suite, right? Data is like air. You don't think
Speaker:about it. Good data engineering is like air. You don't think about it
Speaker:until you don't have any, right? Like, it, you know, if you, you know, it's
Speaker:kind of lost in, in, in the haze, so to speak. And
Speaker:I think that, I think if people,
Speaker:organizations that get their data estates kind of sorted out,
Speaker:are at a competitive advantage to anyone that doesn't, right? And
Speaker:it's very often— that's why I always make a big deal in the intro is
Speaker:that people don't think about data engineering. I've been in hundreds of meetings where
Speaker:the data engineering aspect is often neglected, right? One
Speaker:story in particular was this guy had this great idea for this,
Speaker:that for his organization, and he was going to hire
Speaker:He was going to do all this sort of stuff, integrating all these different data
Speaker:sources, you know, from public, private, enterprise, you name it,
Speaker:right? You name it, he mentioned it, right? It was that type of guy. And
Speaker:he goes, I'm going to need 20 data scientists. And I'm looking at it, I'm
Speaker:like, you're going to need 18
Speaker:data engineers and 2 data scientists, right? Like,
Speaker:you know, like, you know, because like the data engineers, Quote unquote,
Speaker:the old school kind of proper data engineers, I
Speaker:mean, data scientists, they don't want to deal with SQL, right? They just want to
Speaker:get kind of their data. But like, in terms of the sheer
Speaker:amount of systems this guy wanted to integrate and do all the
Speaker:translation, I mean, he's going to need— and he's probably going to need,
Speaker:you know, 50 data engineers realistically, right? Or 20 of them doing
Speaker:the work of 50. But yeah, it just goes to show you
Speaker:that no one appreciates data engineering, right?
Speaker:You know, it's almost invisible. It's certainly invisible when it works
Speaker:well, and it's certainly visible when it
Speaker:doesn't work. But it's also when it's working kind of mediocre,
Speaker:it's also kind of invisible too, right? Yeah, yeah.
Speaker:A lot of times it gets underestimated how much time
Speaker:the work is required to put, I think, right amount of data
Speaker:on the table, right? And that's why a lot of times people don't
Speaker:understand why there is a huge backlog. Hey, I'm asking for simple data, why you
Speaker:are telling me it's too much? So there are 50 people ahead of you and
Speaker:every project is going to take 4 days, right? Takes time to find that data
Speaker:curated and put it somewhere where it can be consumed by, say,
Speaker:analytics use cases or data science use cases.
Speaker:You know, one of the things that comes up a lot is the joke is
Speaker:that DBAs used to stand for don't bother
Speaker:asking, right? But it— but I think that
Speaker:we had an earlier guest a long time ago now basically talk about that
Speaker:one of the important shifts, I think AI really puts the acceleration
Speaker:on the shift of mentality, is that data
Speaker:DBAs, Data engineering managers have to think of
Speaker:themselves less as gatekeepers and more like shopkeepers,
Speaker:right? They're not the bouncers at the door anymore, right? They're the people
Speaker:behind the counter that say, how can I help you? Bonus points if they can
Speaker:set up a system that is— again, I'll go back to the DMV, right? If
Speaker:you go to the Maryland DMV, I'm sure it's even more so where you
Speaker:live, given it's the Valley, right? There's plenty of self-service kiosks
Speaker:where if you just need a new XYZ, You can just scan
Speaker:your driver's license, enter some information, of course,
Speaker:enter some payment information as well, and you, you know, they'll,
Speaker:they'll print up or it's almost fully self-service. And I
Speaker:think that the data engineering organization in the enterprise in the future
Speaker:is going to look a lot more like that than in the way it's looked
Speaker:like historically. Yeah, a lot of it is going to
Speaker:be self-service also, right? People just come in if it's simple tasks.
Speaker:Hey, as an organization, data organization, will
Speaker:serve it for you. You don't need to bother us, or we don't need to
Speaker:bother you. Like, you can just get what you need. If it's complex stuff,
Speaker:then we will put our minions to work. Otherwise, our minions will work
Speaker:completely independently with you. Yeah. So do you
Speaker:have— do you track metrics in your product that says,
Speaker:hey, you know, this saves X number of hours, or is that something
Speaker:that Will come later? No,
Speaker:absolutely. We track it today. So it tracks what kind of tasks
Speaker:have been done already by AI for you, and we
Speaker:try to make sort of a guesstimate also. If this task was
Speaker:done by you manually without using AI, this is how much
Speaker:timewise it would have costed you. So that's one part of it. Second
Speaker:is we do a bunch of benchmarking. against some other
Speaker:tools out there. For example, if you use— if you try to do this
Speaker:task using just vanilla Claude Code or GitHub Copilot, or
Speaker:if you use some other AI solutions that Snowflake and Databricks has
Speaker:put out also, like Codex Code, or dbt has put out
Speaker:Visit, how much better we do against that also. So
Speaker:we do the benchmarking, we publish those benchmarks. That's great. There are open-source
Speaker:benchmarks around this. One is called Agentic Data Engineering,
Speaker:ADE, and the other one is Data Agent Benchmark, DAB.
Speaker:So we do that testing, we put out those benchmarks, and those keep continuously changing
Speaker:as the space itself is so fast evolving. Wow.
Speaker:It's not a real technology until there's benchmarks, right?
Speaker:Yes. In the data world, in infrastructure, we love our
Speaker:benchmarks, right? So why not do it? It's the easiest way to compare different
Speaker:tools and different options. Well, awesome. Uh,
Speaker:and the website is Ultima— Ultima—
Speaker:I'm sorry, an Ultima just drove by. I'm sorry. Let me
Speaker:help you there. It's called Ultimate, so
Speaker:ultimate.ai, and our open source
Speaker:project is called Ultimate Code, so ultimate
Speaker:and code. Just search on it and you'll find the repository.
Speaker:You should find our website also right there. Excellent. Thank you very
Speaker:much. And I appreciate— I know we had some issues rescheduling, so
Speaker:I appreciate your patience with that. And I'm really looking forward to
Speaker:checking this out. You have my curiosity because there's all these like little ideas that
Speaker:I have in terms of, you know, I have datasets and things like that, personal
Speaker:and otherwise, that yes, I would love to go through and
Speaker:organize those. However, I don't have time. Right?
Speaker:Just as a personal project. I'm telling you, when you start running your own home
Speaker:lab, you really start to feel the pain of what IT organizations
Speaker:feel, right? I have an LLM server, and when it's down, my
Speaker:kids tell me, and it just becomes like this, like, nightmare of
Speaker:situations. So, like, the best way to learn is to do. And
Speaker:I encourage everyone out there to build out a home lab with whatever you got
Speaker:and just start messing around with the AI tools to make things easier.
Speaker:And, uh, with that, I'll let the outro music play.