Skip to content

Tackling Data Engineering Challenges with Autonomous AI Agents

In this episode, Frank La Vigne sits down with Pradnesh Patil, co-founder and CEO of Altima AI, to explore how AI is revolutionizing the world of data engineering. Together, they dive into the challenges of modern data stacks, the explosion of tools and technologies, and how AI-powered agents are transforming the way data teams build and maintain complex systems. From automating data pipeline management to optimizing infrastructure and safeguarding governance, Pradnesh Patil shares insights on the next wave of data engineering and what the future holds for professionals in the field. Whether you’re a veteran data engineer or a student curious about the evolving landscape, this episode offers practical advice, industry trends, and an optimistic look at how embracing AI can supercharge your career and your organization’s capabilities.

Links

Time Stamps

00:00 Evolution of data technology tools

04:11 Managing tool stack complexity

09:09 Discussion on AI guardrails and future

10:55 Managing MCB server outputs

14:32 Streamlining Data Pipeline Efficiency

18:48 Managing and Deleting Memories

21:12 AI security and virtual employees

27:16 AI transforming jobs and automation

30:43 Impact of AI on IT Industry

32:32 Upgrading legacy systems with AI

37:10 Importance of Data Engineering

39:10 Changing role of data managers

42:45 Home lab challenges and data management

Transcript
Speaker:

Something breaks and then you have to jump out of the bed in the middle

Speaker:

of the night to get those things fixed.

Speaker:

Hello and welcome back to Data Driven, the podcast where we explore the emerging

Speaker:

industry that is AI, data science, and of course, none of

Speaker:

it's all possible without data engineering. Now, unfortunately,

Speaker:

my favorite, most favorite data engineer in the world can't make it here

Speaker:

today. And I am actually enjoying the sunny but not too

Speaker:

ridiculously hot sunny day here in the suburbs of Baltimore,

Speaker:

Maryland. But I am excited here because I have Pranesh Patel,

Speaker:

who is the co-founder and CEO of a company that

Speaker:

does— makes data engineering a lot more palatable, it sounds like, Ultima

Speaker:

AI. Welcome to the show. Hey, um,

Speaker:

Frank, thanks for having me here. Super excited to chat with you

Speaker:

today. Yeah, so, so tell me about your company, Ultimate AI. Is, uh,

Speaker:

tell me about what it is and what led you to make it.

Speaker:

So Ultimate AI, we are a startup based out of the San Francisco Bay Area.

Speaker:

What we do is we use AI to

Speaker:

automate a bunch of data engineering tasks for the teams out there

Speaker:

who do data work, right? The tasks like building

Speaker:

ELT pipeline, extract and loading of the data, transforming

Speaker:

data, and not just building those pipelines but maintaining

Speaker:

those as well. It's always nightmares. Something breaks and then you have to

Speaker:

jump out of the bed in the middle of the night to get those things

Speaker:

fixed. So building and maintaining data pipelines or

Speaker:

managing your data infrastructure Like in the modern data

Speaker:

stack today, we have Snowflake, Databricks, BigQuery, those things, right?

Speaker:

We need to spend a bunch of effort to manage that infrastructure as well as

Speaker:

optimize that infrastructure, like making sure you have turned on the right knobs on the

Speaker:

infrastructure side as well as you have optimized those queries and pipelines,

Speaker:

etc. So we work on that as well. The whole idea is how we

Speaker:

can use the agents to automate a bunch of that work. so

Speaker:

we can scale our teams even further. I'm glad

Speaker:

you mentioned— I'm sorry, go ahead. No, you asked

Speaker:

me that, hey, how did it all start? So me and my co-founder,

Speaker:

we have worked in B2B enterprise space for a long time. You must have seen

Speaker:

all those data teams every company has. There's so much work involved. There is a

Speaker:

long backlog. We went through the same experiences, and we started

Speaker:

thinking there must be a better solution so that we can handle that

Speaker:

backlog. At the same time, a few years ago, this whole AI wave started

Speaker:

coming in. We jumped in headfirst and started building

Speaker:

agents first to sort of automate simple tasks like writing

Speaker:

documentation or writing some data quality tests. And through that,

Speaker:

the tool and the product evolved from there. And now

Speaker:

we automate bunch of things that are day-to-day things for

Speaker:

data engineering folks. That's a good point. You brought up a good point

Speaker:

because you mentioned Snowflake, you mentioned Databricks, you mentioned all of these

Speaker:

technology tools that the stack used to be a lot simpler in data,

Speaker:

in the data space, right? You, you know, you were either an Oracle shop

Speaker:

or a SQL Server shop and you had, you know, SQL

Speaker:

Server tooling, which those I'm way more familiar with, but also the

Speaker:

Oracle kind of stack too, right? So it went from being kind of a

Speaker:

The tools may have been limited, yes, but the tools also—

Speaker:

there only wasn't really that much of them, right? You picked one side

Speaker:

and stuck with it, right, for most organizations. And then we

Speaker:

had things like Hadoop and Pig and Hive and all of those things kind of

Speaker:

come out. And now this many decade and a half or so

Speaker:

or more, you can tell by my gray hair, that now we

Speaker:

have just an unlimited assortment, it seems, of these

Speaker:

tools, right? And each one of them has their own quirks, their own settings, as

Speaker:

you said, the knobs and dials. Is that something that your

Speaker:

solution offers a solution to? Yeah,

Speaker:

absolutely. Because as you mentioned, the use

Speaker:

cases exploded, and with that, the tool stack exploded as

Speaker:

well. You can't be exploiting 35 different tools to make sure

Speaker:

everything works perfectly and keep your tabs on everything.

Speaker:

So we help tremendously with that, like all those configurations in your

Speaker:

infrastructure as well as building out things. You can't be expert in so many tools.

Speaker:

And on the other side with AI, what has happened is

Speaker:

even these tools as a software are evolving so fast. In a

Speaker:

year, there are like 100 features come out. How can you keep tab on what's

Speaker:

the latest and greatest and what actually fits in my environment and to my use

Speaker:

cases? And, you know, AI for Rescue there, it can

Speaker:

do a bunch of those things for you. No, absolutely. And

Speaker:

I hadn't really thought about that angle of you have to be an expert.

Speaker:

It's like what happened in the software development world

Speaker:

where you had the notion of a full-stack developer where suddenly,

Speaker:

you know, it went from you would be a Visual Basic developer, right?

Speaker:

Or a developer of, you know, web developer. But

Speaker:

then it became, no, you had to be a full-stack developer. You had to know

Speaker:

the data side. You had to know the CSS. You had to know the HTML

Speaker:

and the JavaScript, right? Yeah. Data, I think, has followed a very similar trajectory

Speaker:

in that regard of no one person can do it all. At least

Speaker:

no one person can do it all without the assistance of some kind of

Speaker:

AI. And is that what your product enables? Like

Speaker:

somebody to like basically be a— in the DC area,

Speaker:

they love the term force multiplier. Is that kind of like what

Speaker:

your tool does? Exactly. Exactly. Because

Speaker:

as I think things exploded, use cases exploded, and even

Speaker:

the use of AI has exploded overall, your data

Speaker:

actually fuels all of it. And what's happening is the data projects

Speaker:

are growing exponentially, but our teams are not growing in that size. And

Speaker:

that's why there is a huge backlog that's happening. So the idea is then,

Speaker:

in your words, how we can have some sort of force multiplier that

Speaker:

can do a bunch of things for me automatically. And I focus on some of

Speaker:

the more complex things or something where AI needs more help

Speaker:

and guidance. So think of this as these bunch of agents

Speaker:

are my minions. They're getting the work done. I have like 30 of those and

Speaker:

I'm giving them directions, course correcting, telling them what to do. And

Speaker:

that's how basically the teams are scaling themselves today. But there are so many challenges

Speaker:

in doing that as well, in that whole process. And our

Speaker:

goal is how we can streamline and make that whole process smooth. So you have

Speaker:

teams of these minions which are getting a lot of work done for you.

Speaker:

Yeah, and I think you brought up a very real pain

Speaker:

point, right? They're not hiring tens of people in the data engineering

Speaker:

teams anymore, right? Data engineers are expected to do, to all

Speaker:

be 10x engineers, right? And so

Speaker:

I like the idea of you calling these AI agents minions, one, 'cause

Speaker:

I think the movies are really cute and I have small kids. But

Speaker:

how do you— what's the governance look like on that? Is that something that you—

Speaker:

how did you address that problem? 'Cause I'm sure that comes up quite a bit.

Speaker:

No, I think you touched on very important area, right? Because as a

Speaker:

human, we usually have general sense of understanding and

Speaker:

what needs to be done and what shouldn't be done. But for agents,

Speaker:

it's a piece of code, it's a machine, right? A lot of times it

Speaker:

doesn't have that, compass of common sense. So for example,

Speaker:

that's why you might have seen situations where agent ran a query and that cost

Speaker:

you thousands of dollars. Oh, and there's so many stories of agent

Speaker:

deleted my sensitive data or replicated it. So whether it's sensitive

Speaker:

data, access controls, cost guardrails,

Speaker:

all of those are very important factors around which

Speaker:

we need to develop a solution. Now we just talked about Lot

Speaker:

of tools in the data stack. Every tool has their access layer, RBAC

Speaker:

layer. Now, how are we going to do this, right? So what we have been

Speaker:

doing is, and that's the reason we launch the, our product as open

Speaker:

source project, we call it Ultimate Core. There is a governance

Speaker:

layer that's built in, which allows people to define these guardrails, rules,

Speaker:

and permissions. Lot of guardrails come inbuilt.

Speaker:

And the beauty of that is this is a layer that sits on top of

Speaker:

these bunch of different tools that you are using in your data stack. and gives

Speaker:

us the common ground around governance. So for example,

Speaker:

cost, it makes sure the agent doesn't spend more money than this

Speaker:

for your specific task. Or for agent

Speaker:

itself, you can create a separate view, separate tables so that they don't

Speaker:

inherit like service user permissions or user permissions directly. Because as

Speaker:

a user, I might have so many permissions, but I don't want my agent to

Speaker:

have the same permissions and use my same credentials, et cetera.

Speaker:

So that part also we have solved pretty well.

Speaker:

Yeah, I think, I think calling them minions works out pretty well too, because, you

Speaker:

know, the minions always— there's a whole sequence in the first movie about the dart

Speaker:

gun. Parents, if you know, you know. But

Speaker:

the minions misheard it and built something else. But

Speaker:

no, I think you're right. Making sure these things have guardrails on them,

Speaker:

I think, is— people are going to learn very quickly what happens if you don't

Speaker:

have guardrails, right? millions of dollars a query, or etc., etc.

Speaker:

Actually, recently in the news, and they haven't disclosed what was the core— what

Speaker:

was the core problem, but apparently AWS had—

Speaker:

was giving out bills that were like orders of magnitude higher.

Speaker:

And I can only— again, I have no inside information,

Speaker:

but I can only imagine that there was probably some kind of rogue AI doing

Speaker:

a math error or doing something crazy like that. What do you think the

Speaker:

future of these— this tool space is going to be? Do you think You know,

Speaker:

there'll be more MCP servers, right? What do you think,

Speaker:

um, what do you think that's going to look like, the ecosystem,

Speaker:

the data ecosystem? So how I

Speaker:

see it is based on what we want to do, the tools

Speaker:

get built, right? And as we sort of expand our

Speaker:

horizons to do more and more things, the tools get expanded also, right?

Speaker:

For example, I think MCP server is a great example. At one point

Speaker:

we realized, hey, we need to sort of feed in all this

Speaker:

information to MCP server, uh, to AI agents, right? And then

Speaker:

MCP servers were built as a solution to it. Now

Speaker:

we are hitting the limits of MCP servers themselves where,

Speaker:

you know, okay, governance— anybody can install any MCP server and I have

Speaker:

no control over who uses which MCP server in the organization. Right. That's

Speaker:

happening. Second is The MCB servers are putting out so

Speaker:

much output that my tokens are getting consumed like candy,

Speaker:

right? So then how do I control it where MCB tool outputs are

Speaker:

limited? And we have built some functionality around that, around context comparison

Speaker:

and tool output curtails, etc., especially for data engineering

Speaker:

tasks. Our third big thing is MCB server is like

Speaker:

API, is going to just pull the data from the other system and bring it

Speaker:

to you. But it is not going to do any intelligent filtering

Speaker:

or connecting those dots together. And then now people are looking at

Speaker:

semantic layers and those things, how I can do it also. So now we are

Speaker:

running into the limitations of NCP servers themselves, and then people are trying

Speaker:

to figure out, hey, what's the next thing we need to do to solve this

Speaker:

problem really well? And then people are talking about context graphs

Speaker:

or context store where all of your information is already there,

Speaker:

curated, filtered, smartly arranged, And then you use MCP

Speaker:

server to pull only right information instead of directly interfacing with the

Speaker:

tool, like for example Salesforce or HubSpot, and just dumping everything into your

Speaker:

AI agent as well. So how I see it is, I think

Speaker:

agents are going to become more and more autonomous, more and more ambient

Speaker:

as well, and the information we are going to feed them

Speaker:

is going to become much more curated as well. And there are so

Speaker:

many facets to this. There is a right information feeding angle around

Speaker:

context. There is a governance angle to it, and now. And now I think you

Speaker:

just touched on it. The one big angle that's coming into the play is cost

Speaker:

as well. There, there have been so many stories coming out where people spend

Speaker:

their entire year's worth of budget in like 3 months. I know some stories

Speaker:

which are not public where people's usage like 30x'd in

Speaker:

6 months and now they're like, oh my God, my AI bill is actually same

Speaker:

as my cloud bill now. I never planned for this.

Speaker:

And what is exactly the ROI that people are trying to measure also?

Speaker:

Yeah. So all these questions are coming up, and as the questions and use cases

Speaker:

come up, I believe we'll have more tooling and better solutions.

Speaker:

Is anything— is there anything in particular that your product addresses

Speaker:

to any of these problems? Yeah, so we talked about the governance

Speaker:

piece of it. On the cost side as well, what we started doing

Speaker:

is one of the features we have in the product is context compaction.

Speaker:

So I was talking about MCP tools. Dumping a lot of data, like

Speaker:

especially data-related MCP tools. So what we do is

Speaker:

we curtail the output correctly. We know all these MCP

Speaker:

servers, etc., so that you are not spending too much of a token cost. And

Speaker:

not just the cost, right? What they call is a context rot. If you

Speaker:

dump in too much information, LLM has too many directions to go in. So we

Speaker:

curtail and do that context compaction. Second part we

Speaker:

have introduced is Specifically for data tasks, we create

Speaker:

memories automatically. So not every time agent is starting from scratch,

Speaker:

and it's very curated for data-related tasks. So in that way,

Speaker:

next time when the agent does the same task, it can tap into those memories

Speaker:

that are shared across the organization and get the task done in

Speaker:

very less number of steps. That's an interesting

Speaker:

point because I noticed that, that was one of the When I started

Speaker:

poking around my OpenClaw instance, right, inside of there, there's the

Speaker:

soul, but there's also kind of this memory type of notion and managing

Speaker:

that memory so the context doesn't have to start from

Speaker:

zero every time. And does that really save on tokens? Like, what's

Speaker:

the rough order of magnitude in terms of what the token save is?

Speaker:

No, the memory will save you tremendously because, for example,

Speaker:

let's take an example of data pipelines. Right? So if your data

Speaker:

pipelines are failing, usually there is a pattern, same issues you're going

Speaker:

to see again and again. Hey, data hasn't landed, that's why this pipeline particularly

Speaker:

fails. Now if that thing gets stored in the memory, your agent

Speaker:

is going to check the first thing is has data already

Speaker:

landed? That will save you other 10 other things

Speaker:

that agent would normally try out before coming to that, right? Boom, your

Speaker:

multiple workflows are saved. Your token cost saved, your time is saved

Speaker:

also. We fixed it very quickly, right? So it helps us

Speaker:

tremendously in that way. And the beauty of that is even the memory

Speaker:

layer cannot be very generic, right? You need to understand

Speaker:

as tasks are happening, what kind of memories are important. It needs

Speaker:

to be curated. There needs to be, say for example, somebody's writing

Speaker:

agents and harness for finance, there needs to be a finance-specific memory. that

Speaker:

will remember finance-related things. Similarly, for data work, what we

Speaker:

have created is data agent-specific memories, which works amazingly

Speaker:

well. And since we are talking about memory, right, memory

Speaker:

is not just limited to one session. So for example, what I'm trying to say

Speaker:

is, now if you have data pipelines, probably you have a team of

Speaker:

engineers maintaining that data pipeline. So for example,

Speaker:

I fixed this data pipeline today, for that data not landing

Speaker:

issue, and it gets saved in my memory. Of course, next time

Speaker:

in the regular systems, I try to fix it. Next time it will read from

Speaker:

the memory stored in my, say, code editor or Cloud Code or something like that,

Speaker:

and I can draw to it. But what about somebody else on my team? They

Speaker:

are also going to work on that pipeline, and they might be on on-call. That's

Speaker:

when the pipeline failed. So the memory layer that we have

Speaker:

built, it's a layer that gets shared across the teams and

Speaker:

organization as well. So then it's the— I call this a

Speaker:

hive-like mind in which you are storing the memory, and anybody can

Speaker:

come in and use that hive-like mind and sort of use

Speaker:

that knowledge to do things better. And the big benefit of this is

Speaker:

even for newer people joining your team, right, they don't have to start from scratch.

Speaker:

All this tribal knowledge is stored in that hive mind, which

Speaker:

we call as a memory layer. I like that because then

Speaker:

your first few rounds of learning it, right, you're

Speaker:

literally onboarding this virtual employee, it sounds like, right? It sounds somewhere between a minion

Speaker:

and a virtual employee that can kind of capture that tribal knowledge,

Speaker:

capture that institutional kind of wisdom. Yeah. And so

Speaker:

it's not so much you're spending tokens, you're kind of investing tokens, right,

Speaker:

for the future and training this virtual employee. And

Speaker:

presumably you'll get that, you'll see dividends later on.

Speaker:

Exactly. And this system works with agentic

Speaker:

frameworks that are out there already. People use, say, Claude Code, GitHub

Speaker:

Copilot, Cursor. The system works with those

Speaker:

already. We don't build LLMs, we don't build—

Speaker:

give agentic frameworks, but we build this harness which makes

Speaker:

these tools extremely suitable or powerful when it

Speaker:

comes to data engineering work.

Speaker:

Interesting. Can— does it learn on its own or

Speaker:

can you edit those memories? Right. So like, what if— what,

Speaker:

let's just say we have a, somebody makes a mistake, right? You obviously

Speaker:

wanna mark that and kind of remove that from the memory. Is that, is that

Speaker:

like, obviously you probably edit, it probably adds its own. And,

Speaker:

and the reason why I mentioned this is because when I dove into my

Speaker:

my OpenCLAWS memory file. I thought it was funny what it read about me.

Speaker:

So if people are watching this, you'll see I'm kind of like winking my eyes.

Speaker:

Apparently allergies are really bad today, which I did not factor that in when

Speaker:

sitting outside. And one of the things it learned about me was I

Speaker:

always ask about the pollen levels for the day, which I think is kind of

Speaker:

funny. It said that, you know, I ask about the weather, I ask about stocks,

Speaker:

and I ask about, you know, AI innovations and pollen report,

Speaker:

right? Does this— but obviously I can go in, I can open up a terminal

Speaker:

and edit it myself. But how does your solution— is it built into

Speaker:

the UI or is it like just a config file? So

Speaker:

as you start using the solution, you install, it will start creating

Speaker:

memories automatically. You can, of course, in your prompt give a

Speaker:

specific instruction. As you give the instructions, those get saved also. But you

Speaker:

can say specifically create a memory also. And through our

Speaker:

MCP, it will get that memory created as well. Now the harder

Speaker:

part usually is what if there are bad memories and you want to change memories

Speaker:

or you want to erase those out? The good news is it comes with that

Speaker:

flash stick that I think I remember it from the movie Men in

Speaker:

Black, right? It comes with that. So you can go in, delete your

Speaker:

memories because as I was telling you, we store these memories in a SaaS. So

Speaker:

in that way, they're shared with your team members, they're shared with the rest

Speaker:

of the organization, and there is a granular control you can do who they get

Speaker:

shared with. But at the same time, if you want to update

Speaker:

it, you can go to the UI and get those updated or

Speaker:

deleted as well. Oh, interesting. So you

Speaker:

have that neuralyzer built in. I think that's what the thing is called in Men

Speaker:

in Black. But so you can go back

Speaker:

and you can remove like, hey, everything I did today was terrible, so don't remember

Speaker:

that. What about

Speaker:

security, right? Obviously, There's a lot of trade, you know,

Speaker:

there's a lot of sensitive information that are going to be floated around in the

Speaker:

data engineering space. How does your solution, how does Ultimate AI

Speaker:

kind of address that? So first and foremost,

Speaker:

we don't look at the data directly. Okay, so it's a

Speaker:

harness that comes in, and this harness people can install locally

Speaker:

as well. So if you're some sensitive industry, maybe healthcare

Speaker:

or something like that, You can use it completely

Speaker:

locally, and the LLM solution that you use in the

Speaker:

background, people can use their LLM solution. As I was talking about, Claude

Speaker:

Core subscription or Codex subscription, they can use

Speaker:

that. They can hook up their own models also. We support, say for example,

Speaker:

OpenRouter, all these different models people can use. Even if they want to use

Speaker:

on-premise LLM, we support that as well. On the other side,

Speaker:

as a company and as a platform, We are SOC 2 certified,

Speaker:

pen tested, a bunch of big Fortune 500 companies

Speaker:

use our product already. So even on that side, I'm

Speaker:

sure we can make people's security teams happy because we have done a bunch of

Speaker:

work around it already. Oh, that's interesting. That's good.

Speaker:

Because I mean, the security conversation

Speaker:

comes up in AI, but I don't think it comes up often enough or early

Speaker:

enough. And Obviously, I think that's going to change as more and

Speaker:

more systems get deployed. And obviously, the Fortune

Speaker:

500 companies obviously also take that into account as

Speaker:

well. What would be your advice to people who are data

Speaker:

engineers who are curious about how do I make— how do I get my own

Speaker:

minions, right? Like, how do I start thinking about— let's roll

Speaker:

that up. How do I start thinking about

Speaker:

my job as a data engineer in terms of

Speaker:

I'm managing a dozen potential virtual employees

Speaker:

as opposed to doing it myself, quote unquote, the old-fashioned way.

Speaker:

Yeah, yeah. No, I think that's how we should start

Speaker:

thinking about it if anybody hasn't started going on that

Speaker:

path. There are a bunch of tools out there, and I

Speaker:

believe it's easy to get lost also because there are just so many tools and

Speaker:

things that are happening. on the AI side of the things. And what I have

Speaker:

seen is people try to retrofit the tools for software engineers to data engineering,

Speaker:

and usually that leaves bad taste in their mouth. And sometimes

Speaker:

I heard, hey, AI doesn't work. It's because you're using the wrong tools for data

Speaker:

engineering most of the times. So you need to use the

Speaker:

harnesses or tools that are specifically done for data engineering.

Speaker:

On our side, we open source Ultimate Code Project, so definitely try it out.

Speaker:

It's MIT licensed, completely open source. But in addition to that, what

Speaker:

we have done is on our website, we have put this

Speaker:

course called AI Data Engineer. So there is a section on our

Speaker:

ultimate.ai website, AI Data Engineer, where we have put

Speaker:

curated articles and we keep spending time to make sure all the updated

Speaker:

information is there. So somebody might be like at day 1, somebody might

Speaker:

already be at day 50. You go through that material and it will coach

Speaker:

you. What are the latest tools and what are the best things you can possibly

Speaker:

use to get these things done? Oh, that's really cool. So you

Speaker:

have built-in training, and you did mention that your product is open source. We'll make

Speaker:

sure that the link to the GitHub repo and as well as the, um,

Speaker:

this onboarding, um, this kind of this ment— what would you call it,

Speaker:

a mentoring tool, an onboarding tool? Like, what would you call

Speaker:

that? I just call it training boot camp. So

Speaker:

join it. Even if you're brand new, there is a bunch of useful

Speaker:

information. Even if you have been using it for a while, you know, world

Speaker:

of AI and technology changes so fast. Definitely it will help you

Speaker:

keep tabs on what are the things changing as well.

Speaker:

Yeah, no, I— it's a, it's a very fast-moving space

Speaker:

and it, it's not gotten slower. It's— I think the pace of innovation is

Speaker:

accelerating for good or for bad. What

Speaker:

What would be your advice to someone? I started asking this in the question, like,

Speaker:

there's a lot of computer science students that are very still

Speaker:

in school and they're very worried about AI taking their jobs. I think you and

Speaker:

I have gotten to the point where AI doesn't really take away your job. I

Speaker:

think it changes the nature of your job.

Speaker:

Yeah. But what would be first? My advice would be stick it

Speaker:

through, kids. But like aside from that, what would be

Speaker:

your advice to particularly, so I think a lot of people

Speaker:

are getting started in data engineering and if they hear that there's AI tools

Speaker:

assisting with that or doing that, they may be a little bit afraid or concerned

Speaker:

about the future here. I think the future looks bright, but that's just my opinion

Speaker:

as a, I wouldn't say I'm a perpetual

Speaker:

optimist, but I've kind of been through that cycle of,

Speaker:

oh, you know, I, AI is taking away my job, but I really kind of

Speaker:

see the nature of it. It just changes the nature of my job.

Speaker:

Yeah, yeah. If I can summarize this, I have heard those

Speaker:

things, hey, is AI going to take my job? And I always tell

Speaker:

if anybody says this to me, in my opinion,

Speaker:

it's not AI that's going to take your job, but it's going to be somebody

Speaker:

who can use AI is going to take your job. So if you don't upscale

Speaker:

yourself, because I completely resonate with your point. Our

Speaker:

jobs and roles are changing rapidly in this new world, so

Speaker:

upskilling yourself for AI is very, very important. And being

Speaker:

on that cutting edge of the AI and data engineering and all that data

Speaker:

work, that's the most important part. How I'm feeling

Speaker:

is going to unfold is like how 100, almost 100 years

Speaker:

ago, Industrial Revolution unfolded, right? There were, for example,

Speaker:

there were people who are making clothes by hand, right?

Speaker:

Then the machines came and we got operators who would operate those machines

Speaker:

to produce clothes at a humongous scale and see how

Speaker:

many clothes we have today, right? Same thing is going to happen

Speaker:

with AI in general. It's going to give us those machines

Speaker:

which will increase the overall throughput of all the people

Speaker:

out there for different professions. It's going to bring in a lot of prosperity.

Speaker:

But at that time, people who were making clothes by hand, they went out of

Speaker:

job. Maybe some of them learned how to operate the machines, but that's

Speaker:

the important part. Learn to operate the machines. Like, don't just say that, hey,

Speaker:

AI is going to take my jobs. Learn to use AI, embrace it. Your job

Speaker:

is not going anywhere there. That's a really good way to put it.

Speaker:

One of the examples I like is, um, there's self-driving cars,

Speaker:

right? But there's also a lot of cars have adaptive cruise

Speaker:

control. So for those not familiar with it, you know, it's basically cruise

Speaker:

control, which will keep your speed, but it also has proximity sensors. So

Speaker:

it'll apply the brake, right? And kind of keep you within a certain distance of

Speaker:

the car ahead of you, et cetera, et cetera, et cetera. And also can keep

Speaker:

you inside your lane. When I got that in, in

Speaker:

my car for the first time, I felt like— It's a good analogy for AI

Speaker:

because I'm still controlling the car. Right. I'm not fancy enough

Speaker:

to have a full-on, you know, self-driving car, but, um, but I

Speaker:

did notice that I became more like a captain of a ship, so to speak,

Speaker:

where I would— I felt like I was guiding the car which direction to

Speaker:

go. Now, in that case, I'm not a big fan of driving around

Speaker:

all the time, but so it would be nice to have a full driving car.

Speaker:

But I think that's kind of a good analogy for how,

Speaker:

how jobs are going to look like, right? And if you're a software engineer, you're

Speaker:

not going to be writing every line of code by hand anymore. I think,

Speaker:

to your example, I think, you know, when humans were sewing, doing the

Speaker:

sewing, right, individual stitch by stitch, clothes were made. And accordingly,

Speaker:

clothes were expensive. And I think we're going to see that kind of happen with

Speaker:

AI, right? Whether it's code, whether it's data engineering tasks.

Speaker:

Because certainly I think anyone out there who is in a professional

Speaker:

position where they're doing data engineering, There's plenty of

Speaker:

things that are just kind of on the back burner that they would do to

Speaker:

make their jobs more efficient, right? And I just think that,

Speaker:

I think by having AI enabling,

Speaker:

you have a lot more cognitive free room. And all those things

Speaker:

are on the back burner, on the whiteboard somewhere, or kind of on a Post-it

Speaker:

note somewhere about, hey, you know, I bet I can improve this process by

Speaker:

X number of percent, but they're too busy keeping the machines

Speaker:

working, right, keeping everything working. So with AI, I think you can have a

Speaker:

bit of cognitive surplus where you can tackle those. Sorry, I cut you off. You

Speaker:

were about to say something. No, nothing to add there. I think

Speaker:

that's so spot on. This is where the world is going.

Speaker:

That's cool. That's cool. So what's next for your company? Like, what,

Speaker:

what's, like, what other big challenges do you think you

Speaker:

all are going to take on? So one thing is, of course,

Speaker:

right now we have been automating few of the data engineering use cases,

Speaker:

but we are planning to expand to more and more use cases and at

Speaker:

the same time improve the support of the data stacks

Speaker:

we support as well. For example, you're talking about SQL Server,

Speaker:

maybe older Hadoop, Hive stacks. There's so much information

Speaker:

out there, and being early-stage company, we support certain

Speaker:

stacks, certain stacks we don't support. So the plan is to expand

Speaker:

into more ecosystems and at the same time expand for more

Speaker:

use cases as well. Yeah, that's cool. Yeah, and it's funny

Speaker:

because I know there's a lot of Hadoop legacy solutions out

Speaker:

there. Somebody told me, and I don't want to pick on Hadoop, right? But

Speaker:

Hadoop was one of the first, you know, big data solutions out there. And

Speaker:

I would imagine there's a lot of— somebody told me that it was moved to

Speaker:

the— what's Apache call it? The attic? the attic or the basement where it's basically

Speaker:

considered a legacy product now. And, you know, obviously

Speaker:

there's a lot of engineering— data engineering projects historically

Speaker:

have not turned over that quickly, right? Like, you know, there were— there are probably

Speaker:tch jobs I wrote in the early:Speaker:

And that's not that unusual. So I think that

Speaker:

because data modernization projects even if they're

Speaker:

needed, they may not happen because people are too busy keeping

Speaker:

the machines running, right? So if you can kind of— I think people are missing

Speaker:

the point about AI taking their jobs. I know we're back on this again. You

Speaker:

can always tell when the coffee hits me, right? Like, my mind goes in like

Speaker:

3 different places at once. But the release

Speaker:

of cognitive service, cognitive surplus,

Speaker:

I think is going to do a lot of good things for The IT

Speaker:

industry, and ultimately, I think, for the broader economy as a whole, right?

Speaker:

I think you, you mentioned prosperity before, and I know it's a, a lot of

Speaker:

people are very anxious about what AI can do. And you live, you know,

Speaker:

you're in the Bay Area, you're probably at the epicenter of a lot of this

Speaker:

back and forth, right? But I think that

Speaker:

by releasing some of this, the human minds

Speaker:

from the drudgery of some of the, the machinery that we have today,

Speaker:

I think has enormous potential for improving IT, right?

Speaker:

You know, take any kind of like regular person out there who, who has to

Speaker:

interact with their IT department. Would they describe it as a wonderful experience, right?

Speaker:

Would they describe it as a, you know, or is it more like going to

Speaker:

the DMV, right? Probably more like the DMV. Although

Speaker:

I will say credit where credit is due. In Maryland, I had to spend— I

Speaker:

recently got a new car and it wasn't as bad as I feared

Speaker:

it to be. You know, maybe they're using AI, I don't

Speaker:

know. Yeah, yeah. And as you mentioned, right, a lot

Speaker:

of this work we sort of parked it for later. Like you were talking

Speaker:

about batch jobs or Hadoop jobs, etc. There's still a bunch of those things running

Speaker:

because migrations had been nightmares, right?

Speaker:

Like people spend millions for years. But now I think, now

Speaker:

if I look at it with the angle of AI, AI can automate

Speaker:

90-95% of those migrations as well. Now, who knows, like, even

Speaker:

that whole modernization and digitization

Speaker:

in the data world, that will start accelerating much faster because of

Speaker:

this. And a lot faster, different companies can be on the

Speaker:

most modern technology. And they can, they can take advantage of

Speaker:

the newer features, the optimization technology. Yeah, I mean,

Speaker:

because no, no company in their right mind is going to suddenly, you know, say

Speaker:

we're going to hire, you know, we have this system, we can either

Speaker:

A, keep throwing a little bit of money at

Speaker:

it per year, right? And keep it going indefinitely. Or

Speaker:

B, we're going to hire 100 new data engineers and all these project

Speaker:

managers and things like that. We're going to spend billions on upgrading it. They're not

Speaker:

going to do that, right? But they do need to upgrade, right? So then if

Speaker:

it becomes a, well, you know, if you just spend just a little bit

Speaker:

more than what you do have to do to keep the lights on, and you

Speaker:

kind of have AI take a lot of the brunt work of that, right? Instead

Speaker:

of hiring 400 people, you can probably hire, you know, say

Speaker:

40, right, to do the migration. It starts to become more

Speaker:

palatable, right? It starts to become more justifiable in a, in

Speaker:

a balance sheet. And, and, you know, I think a lot of the breaches

Speaker:

we've seen have been from improperly

Speaker:

configured systems, right? Or legacy software

Speaker:

that hasn't been patched or legacy software that has no business still running

Speaker:

in this day and age. And I think you're right. I think AI is really—

Speaker:

I think if you put the right type of glasses on, AI

Speaker:

is a way to solve the problems that we've just

Speaker:

been pushing back for years and years. And technical

Speaker:

debt, that's the word I'm looking for. I think AI has the potential to

Speaker:

be a great way to pay down technical debt. Yeah.

Speaker:

But So while we're talking about

Speaker:

that, like, what is the thing that your customers who are successful on your

Speaker:

platform, what is the thing that they're most happy with at the end

Speaker:

of the day? What are they like, they call you up and they say, wow,

Speaker:

Pranesh, I'm so glad we did this because

Speaker:

what is it, increased productivity? Is it something else?

Speaker:

A few things, right? So for example, one example I gave was

Speaker:

we can manage people's infrastructure using AI. Right? And

Speaker:

we can optimize it. Now AI does it at scale

Speaker:

at which humans can't even do it. For example, it can change the configurations

Speaker:

300 times in a day. We can't do it, right? Like, how can we

Speaker:

change the configuration of single machine 300 times a day? You need 150

Speaker:

people doing just that. Yeah. And

Speaker:

especially in the space of data, your workloads are always changing, right?

Speaker:

As your data changes, your workload changes. So one of the things we have done

Speaker:

is that infrastructure optimization, where we have built agents which can

Speaker:

automatically analyze your infrastructure, change the configurations. We have built agents

Speaker:

which can analyze the data pipelines and automatically optimize them. And

Speaker:

of course, that saves tons of time for engineers, but at the same

Speaker:

time, your infrastructure runs at 90 to 100%

Speaker:

utilization, and those pipelines are always top-notch when it comes to

Speaker:

optimization because people run millions of SQL queries and hundreds of thousands of

Speaker:

pipelines. Now AI does that. The result is not just the

Speaker:

engineering hours saving, but real dollar savings on

Speaker:

the infrastructure cost also. So some of our biggest customers, we

Speaker:

have saved them millions of dollars in their Snowflake and

Speaker:

Databricks environment by optimizing this.

Speaker:

And yeah, that's super happy about it. They're like, hey, tool pays for

Speaker:

itself multiple times over. We are getting all these engineering productivity also. We're saving

Speaker:

tons of engineering hours. But hey, right there and then you're saving me

Speaker:

infrastructure costs also by automating that process by a huge margin.

Speaker:

No, it's a great way to— that's a great way to look at it. And

Speaker:

I really think there's so much opportunity. I know there's a lot of people, there

Speaker:

are a lot of naysayers now about what AI looks like and,

Speaker:

and, you know, will we, will we realize the gains that were been promised?

Speaker:

I think we will. I think it's going to be in places where we may

Speaker:

not think of, right? No one You know, and data engineers will be the first

Speaker:

people to tell you, like, they're not usually first of mind for a lot of

Speaker:

people in the C-suite, right? Data is like air. You don't think

Speaker:

about it. Good data engineering is like air. You don't think about it

Speaker:

until you don't have any, right? Like, it, you know, if you, you know, it's

Speaker:

kind of lost in, in, in the haze, so to speak. And

Speaker:

I think that, I think if people,

Speaker:

organizations that get their data estates kind of sorted out,

Speaker:

are at a competitive advantage to anyone that doesn't, right? And

Speaker:

it's very often— that's why I always make a big deal in the intro is

Speaker:

that people don't think about data engineering. I've been in hundreds of meetings where

Speaker:

the data engineering aspect is often neglected, right? One

Speaker:

story in particular was this guy had this great idea for this,

Speaker:

that for his organization, and he was going to hire

Speaker:

He was going to do all this sort of stuff, integrating all these different data

Speaker:

sources, you know, from public, private, enterprise, you name it,

Speaker:

right? You name it, he mentioned it, right? It was that type of guy. And

Speaker:

he goes, I'm going to need 20 data scientists. And I'm looking at it, I'm

Speaker:

like, you're going to need 18

Speaker:

data engineers and 2 data scientists, right? Like,

Speaker:

you know, like, you know, because like the data engineers, Quote unquote,

Speaker:

the old school kind of proper data engineers, I

Speaker:

mean, data scientists, they don't want to deal with SQL, right? They just want to

Speaker:

get kind of their data. But like, in terms of the sheer

Speaker:

amount of systems this guy wanted to integrate and do all the

Speaker:

translation, I mean, he's going to need— and he's probably going to need,

Speaker:

you know, 50 data engineers realistically, right? Or 20 of them doing

Speaker:

the work of 50. But yeah, it just goes to show you

Speaker:

that no one appreciates data engineering, right?

Speaker:

You know, it's almost invisible. It's certainly invisible when it works

Speaker:

well, and it's certainly visible when it

Speaker:

doesn't work. But it's also when it's working kind of mediocre,

Speaker:

it's also kind of invisible too, right? Yeah, yeah.

Speaker:

A lot of times it gets underestimated how much time

Speaker:

the work is required to put, I think, right amount of data

Speaker:

on the table, right? And that's why a lot of times people don't

Speaker:

understand why there is a huge backlog. Hey, I'm asking for simple data, why you

Speaker:

are telling me it's too much? So there are 50 people ahead of you and

Speaker:

every project is going to take 4 days, right? Takes time to find that data

Speaker:

curated and put it somewhere where it can be consumed by, say,

Speaker:

analytics use cases or data science use cases.

Speaker:

You know, one of the things that comes up a lot is the joke is

Speaker:

that DBAs used to stand for don't bother

Speaker:

asking, right? But it— but I think that

Speaker:

we had an earlier guest a long time ago now basically talk about that

Speaker:

one of the important shifts, I think AI really puts the acceleration

Speaker:

on the shift of mentality, is that data

Speaker:

DBAs, Data engineering managers have to think of

Speaker:

themselves less as gatekeepers and more like shopkeepers,

Speaker:

right? They're not the bouncers at the door anymore, right? They're the people

Speaker:

behind the counter that say, how can I help you? Bonus points if they can

Speaker:

set up a system that is— again, I'll go back to the DMV, right? If

Speaker:

you go to the Maryland DMV, I'm sure it's even more so where you

Speaker:

live, given it's the Valley, right? There's plenty of self-service kiosks

Speaker:

where if you just need a new XYZ, You can just scan

Speaker:

your driver's license, enter some information, of course,

Speaker:

enter some payment information as well, and you, you know, they'll,

Speaker:

they'll print up or it's almost fully self-service. And I

Speaker:

think that the data engineering organization in the enterprise in the future

Speaker:

is going to look a lot more like that than in the way it's looked

Speaker:

like historically. Yeah, a lot of it is going to

Speaker:

be self-service also, right? People just come in if it's simple tasks.

Speaker:

Hey, as an organization, data organization, will

Speaker:

serve it for you. You don't need to bother us, or we don't need to

Speaker:

bother you. Like, you can just get what you need. If it's complex stuff,

Speaker:

then we will put our minions to work. Otherwise, our minions will work

Speaker:

completely independently with you. Yeah. So do you

Speaker:

have— do you track metrics in your product that says,

Speaker:

hey, you know, this saves X number of hours, or is that something

Speaker:

that Will come later? No,

Speaker:

absolutely. We track it today. So it tracks what kind of tasks

Speaker:

have been done already by AI for you, and we

Speaker:

try to make sort of a guesstimate also. If this task was

Speaker:

done by you manually without using AI, this is how much

Speaker:

timewise it would have costed you. So that's one part of it. Second

Speaker:

is we do a bunch of benchmarking. against some other

Speaker:

tools out there. For example, if you use— if you try to do this

Speaker:

task using just vanilla Claude Code or GitHub Copilot, or

Speaker:

if you use some other AI solutions that Snowflake and Databricks has

Speaker:

put out also, like Codex Code, or dbt has put out

Speaker:

Visit, how much better we do against that also. So

Speaker:

we do the benchmarking, we publish those benchmarks. That's great. There are open-source

Speaker:

benchmarks around this. One is called Agentic Data Engineering,

Speaker:

ADE, and the other one is Data Agent Benchmark, DAB.

Speaker:

So we do that testing, we put out those benchmarks, and those keep continuously changing

Speaker:

as the space itself is so fast evolving. Wow.

Speaker:

It's not a real technology until there's benchmarks, right?

Speaker:

Yes. In the data world, in infrastructure, we love our

Speaker:

benchmarks, right? So why not do it? It's the easiest way to compare different

Speaker:

tools and different options. Well, awesome. Uh,

Speaker:

and the website is Ultima— Ultima—

Speaker:

I'm sorry, an Ultima just drove by. I'm sorry. Let me

Speaker:

help you there. It's called Ultimate, so

Speaker:

ultimate.ai, and our open source

Speaker:

project is called Ultimate Code, so ultimate

Speaker:

and code. Just search on it and you'll find the repository.

Speaker:

You should find our website also right there. Excellent. Thank you very

Speaker:

much. And I appreciate— I know we had some issues rescheduling, so

Speaker:

I appreciate your patience with that. And I'm really looking forward to

Speaker:

checking this out. You have my curiosity because there's all these like little ideas that

Speaker:

I have in terms of, you know, I have datasets and things like that, personal

Speaker:

and otherwise, that yes, I would love to go through and

Speaker:

organize those. However, I don't have time. Right?

Speaker:

Just as a personal project. I'm telling you, when you start running your own home

Speaker:

lab, you really start to feel the pain of what IT organizations

Speaker:

feel, right? I have an LLM server, and when it's down, my

Speaker:

kids tell me, and it just becomes like this, like, nightmare of

Speaker:

situations. So, like, the best way to learn is to do. And

Speaker:

I encourage everyone out there to build out a home lab with whatever you got

Speaker:

and just start messing around with the AI tools to make things easier.

Speaker:

And, uh, with that, I'll let the outro music play.