Your AI Agent Doesn't Know Your Business | Context Layers Explained
By MotherDuck
Summary
Topics Covered
- Your data definitions will outlive your career
- When your LLM 'infers,' it is guessing
- Open standards die before anyone adopts them
- Skills fail because humans forget to use them
- Tell your LLM to say 'I don't know'
Full Transcript
So, um, thank you very much, Jim, for your presentation because you kind of laid the groundwork for everything we're going to be talking about. And my intent
today is to engage all of you in your and have you participate and have you tell us your frustrations around all of these problems and we're going to talk about some of the solutions people are
coming up with.
Are those the right solutions?
We don't know. You know, this is like a brave new world that we've all we're all heading into, but I will tell you this much.
This problem has existed since the '80s and I've been around a long time and I've heard them talk about this absolutely forever.
Oh, not close enough. Thank you.
Um So, back in the 1980s when IBM was first, you know, coming out and and the idea of a personal database was starting to arise, people started talking about
data. And how are you going to define
data. And how are you going to define your data and how is it going to get used?
And then in the 1990s we came up with data warehousing and it became even more important because now we were collecting everything.
Like every single click, every single anything that was going through people were collecting.
And then in the 2000s it turned into data governance. Gartner Group called
data governance. Gartner Group called data governance the single most important aspect of computing for the the century. Yeah, well, that died.
Then we got into 2010s and we got into the lakes and the idea of cataloging your data and now we're in the 2020s and we're still talking about it.
So, my prediction, having been around for a while, is we'll be talking about this when I'm long gone. Data is hard.
Defining your data is hard. It's a hard process.
So, I want to start with a little terminology because I see these two concepts being conflated quite a bit.
There's a semantic layer and there's a context layer.
And semantic layers are pretty cut and dry. You know, they're your schema,
dry. You know, they're your schema, they're your joins, they're your tick column annotations. It's all that sort
column annotations. It's all that sort of stuff.
There is a lot of people here I know that are in working at companies that are producing semantic layer products.
So, you guys know your product better than I do, but Clyde and I had lovely sessions talking about these and this is my definition that I've come up
with. Do people agree with that
with. Do people agree with that definition for a semantic layer?
Is that generally how you think about it?
Does somebody have a different definition they'd like to share?
Okay, so I'm going to go with that one.
Your context layer is a whole different animal. Your context layer is your
animal. Your context layer is your definitions of business terms. It documents specific business knowledge.
Think about when you onboard a new employee into your company. And they're
sitting down to write their first set of reports for the CEO, for the board meeting and they're like, "Oh, no, no, no. Yeah, don't include that guy cuz
no. Yeah, don't include that guy cuz that company has this problem. No, no,
no, don't do this over here." Right?
"Oh, well, we have an exception for that." Right? Everybody has it.
that." Right? Everybody has it.
Where do you codify that? Where do you put that knowledge in? So, if somebody comes in in a natural language query, they don't have the guy standing next to him, "No, no, no, no, no. We don't
define it for them cuz he plays golf with the other guy, CEO, and they've got this like special deal going." Like,
where do you put all that?
That's your context layer. That's your
color. That's really explaining your data.
Does anybody disagree with that?
Okay.
[snorts] So, this is Clyde did this um cool little graph for me um comparing these. So, for context, he called it
these. So, for context, he called it docs and slacks and human corrections and past exceptions.
Past exceptions is a big one.
Defining terms. So, what if you're a home builder?
Does the word model mean something different to you than it means for anybody else?
Sure.
If an LLM doesn't know that, what's it going to do?
If you don't give it your business terms, if you don't define what revenue means to your LLM, what's it going to do?
I'm curious. What do you think it does?
Cuz I asked Claude and Chat and Gemini.
It guesses.
They don't call it guessing.
What word do you think they use?
I don't know.
It's my favorite. They infer.
So, I looked it up.
[laughter] [gasps] Under According to the Cambridge Dictionary, infer, in fact, means to guess. They have no other option. You
guess. They have no other option. You
haven't given them any guidelines for what those terms mean to you. Think
about the word class.
Oh goodness, how many different definitions can you have for the word class? And if it means something in your
class? And if it means something in your business, you should probably put the right one in so it knows how to define this.
So, right now there is the semantic Everybody's kind of buzzing about all of this. I've been going to a ton of
this. I've been going to a ton of meetups here in San Francisco and Databricks and Summit Snowflake Summit and all the rest of it and AI Council and HumanX, all of them had
presentations on this. And there's
actually a guy who's got a podcast called "WTF is a Semantic Layer?" And
it's pretty good. I really like recommend listening to it because he invites a lot of people on the podcast that are talking about this and writing products around it, which I think is the
most important thing for most of us.
So um they have tried now to come up with something called the open semantic interchange.
Right? What's the problem with these standards?
Anybody want to venture a guess what happens with the standards as they come out?
What is the biggest problem?
No one agrees on anything.
That's the first one.
They seem to be old by the time they come out.
They seem to be old by the time they come out.
There's lots of them.
There's lots of them.
What's the other biggest one that Gartner always um exposes?
Adoption.
Sure, you can have a standard. You have
a bunch of very bright people sitting around talking about this.
[snorts] But who adopts? How fast do they adopt?
And you're right, by the time that you get a kind of a groundswell, it's old.
It's stale. It needs to be rethought.
So, we're back to square one. And so,
this has been going on for a very long time.
So, they did specialize they did finalize a specification Q1 2026.
Anybody here know what it is?
Yes.
What is it? Do you
It's on the spec.
Yeah, do you know what the spec is?
It's on GitHub, yeah.
How many people have listened to it, incorporated it to your workflows, it's now becoming part of your business?
Yeah. So, you see my point.
Okay.
It has some support across 50-plus platforms, but that's support. That's
Oh, yeah, that sounds like a great idea.
Yeah, we should really do that.
And then the vendors go off and build their own product with their own specifications.
And now we're back to square one.
There's no uniform agreement.
So, vendor specific solutions, um, they argue that it's better to use a vendor supplied semantic product because it's tightly coupled
with their solutions. It takes advantage of all the infrastructure of the product. You know, it has a lot of
product. You know, it has a lot of pluses.
What happens when you want to leave?
What happens when you don't want to use that product anymore?
Now, everything is tightly coupled with that product and you're now vendor dependent.
How many of you have embraced the idea of iceberg?
You like the fact that it's vendor independent?
Open format?
All that good stuff?
Well, this is the opposite. So, I hear people are going, "Oh, yeah, no iceberg, really great idea. Oh, and by the way, I'm I'm tying all my semantic stuff to this, you know, ETL tool I'm using." And I kind of look
at them going, "Okay.
Kind of interesting."
The pro for the semantic interchange is it separates what a metric means from where it is run. It makes it like iceberg. Iceberg is completely isolated.
iceberg. Iceberg is completely isolated.
It's It's just a format. It's a catalog.
Anybody can talk to it. You can leave when you want.
So, that is something that they've modeled the OSI after. And the cons are you can't force it.
You simply can't get people to adopt it.
Vendor specific solutions, their pros are governance depth, sustained engineering investment. All these things
engineering investment. All these things are true. But what if you get tied up
are true. But what if you get tied up and all of a sudden you're going, "My spend is what per month?"
I can't This is unsustainable. I have to find a less expensive option. Well, now
you're in a shift well, and shift scenario and that is very expensive to get out of.
Anybody have any other alternatives?
What's going on in your organizations?
This is really I really would like people to um give us their ideas.
You guys had too much pizza. You should
have waited. We should have made you hungry.
The way I think about it is I just plan for a mess. I don't count on standards.
It will be messy, and I have to make sure I do things that survive despite the messiness.
Interesting.
Yeah.
What do you do to survive the messiness?
Um use a lot of AI, right?
Yeah, I mean, the way we handle our we have like, a living version ontology.
We'll we'll kind of manage it. It's kind
of pretty about the key thing we're doing a lot of other stuff but our primary one thing.
Anybody else?
Okay.
The context layer. How do you codify institutional knowledge?
Like, in pretty much every large company I've ever been in, there's always one guy in the corner cube that's been there forever.
He knows where all the bodies are buried. He knows all the reasons why
buried. He knows all the reasons why everything happened the way it happened.
How do you get a brain dump of that person into your context layers so that institutional knowledge is available to your LLMs when they come through?
Really hard to do. Very hard to do.
It takes stepping back and thinking about what your business actually relies on to get the right answers.
So, there's rag pipelines. How many
people using rag pipelines?
Quite a few. How are they working out?
Good bad frustrating?
Good.
Okay.
Claude mentioned that they can't reason across time and relationships. Do you
find that to be true, those of you that are using them?
Not really.
We do use that I find for product documentation.
Okay.
As long as we have different versions laid out we are finding that it works pretty good.
Okay.
Can you repeat that comment?
Yeah, can you repeat it for Oh, yeah. Um here.
Oh, yeah. Um here.
Uh we use it for uh product documentation as long as the documentation has versioning built in.
Uh it does produce pretty good results.
Anybody else? Anybody else? Who's using
rag?
You said you like yours.
Yeah, it's working pretty well. Um I
don't know. I found it to be fine personally, but I think we also try to document it really well like excessive I guess commenting or documentation,
stuff like that. But also my startup is at a pretty early stage, so we don't have like many decades of different documentation, things like that.
How about uh skills?
Skills in your GitHubs.
How many of people are doing skills?
I see a lot of head nodding. Is that
working out well?
Okay, so I have a question for you about skills.
Um what is the one thing humans are notoriously bad at? Like notoriously bad at.
Prioritization.
Okay, he says prioritization. Anybody
have another one?
Consistency.
Consistency. Anything else?
Time estimates.
Time estimates.
Not lying.
Not lying, yeah.
Uh putting things in writing.
Putting things in writing.
Multitasking.
Multitasking.
Naming things.
Naming things.
I've got a big one for you. When was the last time you went to the grocery store and you had a mental list of everything you needed?
And you get up you go to the grocery store, you check out, you're in your car, you're in your driveway, you're in your garage, and you're like, "Oh, I forgot that."
Some ingredient for the big dinner party you're having and your partner is looking at you going, "You forgot that.
You have to go back."
Humans aren't good at remembering. We
just aren't. And and we really don't even know why. Neuroscience hasn't
figured it all out yet, right?
So, my concern with skills is you have to remember to use them.
Like there's a data warehouse that we compete with and I'll not name them other than it's really cold. Anyway, um
they have in their new horizon context layer, you have a slash, and then you have to hunt for the skill that you want the LLM to use.
And I'm like, "That ain't going to work." I can tell you right now that
work." I can tell you right now that ain't going to work, right? So, you have to have a way that forces your LLM to use those skills. You can write skills,
all for them, cuz it's a really good way of documenting that institutional knowledge, but how are you going to get people to remember to use them?
There are also dedicated context layer platforms coming out. A lot of people are are a lot of startups are working on this prob- problem, coming out with products to solve this problem.
But it is a problem. So,
a problem. So, MotherDuck.
Now I get to talk about what how we're approaching this.
So, um MotherDuck is very, very interested in what people think and what problems people are encountering. And we
really like hearing from people who are frustrated because that helps us build a better product. And because we are a
better product. And because we are a startup, we're a little nimble.
So, we can really take user feedback and make very good use of it. So, we did. We
listened to what a lot of people were saying about context layers and what they needed.
And they started building using customers frustrations as as a starting point. So, Hamilton Euler, who's our
point. So, Hamilton Euler, who's our head of um UI and is working on this. Um
had this great quote, and so I put it in. What we're building is a context
in. What we're building is a context layer that lives alongside the tables in Motherduck, so the right knowledge gets surfaced automatically
when an agent touches the relevant data and can be updated both implicitly and explicitly.
So, this product is being released imminently.
You have a question? Oh, no.
Well, this product is being released imminently, um and I think I'm allowed to to tell what it's called.
Neris? [snorts]
Um Yes.
Ducklings?
Hm?
Ducklings?
No. We We do ducklings for our instant sizes. Had to come up with something
sizes. Had to come up with something new.
What's he doing?
Find the relevant data and And what is he doing for them?
Going on the There we go. They're called guides.
It's a way to guide your LLM through your your institutional knowledge. So, here's
what we came up with.
And again, I'd love to hear people's reactions to this.
It is B1 is coming out. But, you know, lots of room for improvement on B2, should it be necessary. So, it's stored natively in
necessary. So, it's stored natively in Motherduck.
You don't have to have You can have it in a GitHub. You can have it loaded from a GitHub, but it's stored natively in your database. You can do a guide per
your database. You can do a guide per database. So, think about that. If
database. So, think about that. If
you're a customer-facing analytics sort of company and you have databases per customer because each one of your customers has their own little wiggies,
you can put a guide per customer in that database that gets read by any natural language query coming through.
Um let's see.
Uh you can create, update, move, and change visibility from SQL.
Uh you can have your agents build it for you, you know, all that usual stuff.
But what I like the best is personal and org scoping. So, you can do an org
org scoping. So, you can do an org scoping level where your admins and your C-suite and all those people have gotten together and made all, you know, said this the stuff that's important to us.
But maybe you have something like some terminology that you use when you're thinking about the data. You can have personal scoping.
And one person that spoke to me said, "Wow, that'd be kind of interesting. So,
I could as an admin, I could watch what they're putting in for personal scoping and then decide whether that's something we should adopt as an organization."
And I went, "Huh, that's a really good idea."
idea." The other way you could use that is if you're going to let an agent start trying to learn your data and and making comments on what it sees,
the trends and the usage by looking at the queries that are coming through, it can put its own scoping in personal.
Won't touch your orgs. It'll be
impersonal.
So, that allows you to review what they've done.
I just have a quick random survey. How
many of you would trust an LLM to write a context layer for you?
It should start.
[laughter] I like that answer. It should start. I
agree.
Is it going to be 100% accurate?
Will humans be?
True. Will humans be? That's a good good answer.
But it it you can give it a place to start thinking about it. And you can give it a place to start and then you can train it. You can tell it, "Hey,
you're spot on there. No, missed on that one." There will also be versioning.
one." There will also be versioning.
You can keep this in your GitHub.
So you can do all the lovely GitHub version control.
And there will be it will be embedded in our UI as well.
So this is just an idea of what version one is going to look like as it's coming out.
And this is what an example guide would look like.
You can put the actual SQL in to your guide.
A lot of people like doing that. They
they find a whittle down that SQL statement. It's exactly what they want.
statement. It's exactly what they want.
And then what happens when your schema changes? Huh. Yeah.
changes? Huh. Yeah.
Can everybody see this?
I should probably walk through it. So
there's a term and a meeting.
Like, you know, what does churn mean? What does
contraction mean? What does expansion means? What's a quarter?
means? What's a quarter?
If you're not on the calendar quarter and you tell Claude to go off and give you the quarterly results, what's he going to do?
January February March April May June. And you come back and you're like,
June. And you come back and you're like, "Oops.
Forgot to tell you that's not our our fiscal year."
fiscal year." If you can actually put in gotchas.
Hey, you know, for this customer, remember they play golf.
[laughter] He gets an automatic discount.
You can tell it exactly where the data lives and you can give it the SQL statement.
So, now you have a place to kind of wrap up everything and grow it. This is
how an agent would use that context layer.
The user comes in, asks the natural language question.
The agent goes to look at the get guide.
It looks for guides and D on Motherduck MCP server, and it receives the tree of existing guides and custom instructions.
Now, it doesn't have to infer or guess. So,
or guess. So, my question to you, does this look something like something that's useful?
I see some nods. Do I see any of these?
Cuz if I do, I want you to tell us why.
Yes sir.
Um I think it looks great, for sure, but as with everything, it feels like it definitely encourages less like hands-on behavior, right? And then my question
behavior, right? And then my question is, how do you stop it from having the same like schema drift problem? And how
do you stop one inaccuracy from compounding and then becoming a structural issue that spreads?
Yeah, that's a really good question. Um
you would have to, obviously, write a new version of guide.
So, if your schema does start to drift, you would have to be aware of that. That
is something I would actually We have another product called dives. I would
actually write a dive to find out every time a schema changes, and that would notify me that I need to adjust my guide.
So, I I don't know how many of you are aware of the the end-to-end uh solutions that Motherduck has been coming up with for agentic workflows. Uh guides is I
call getting your ducks in a row. Now,
we have flights, which is how to get your data from point A to point B. And
last is dives, which is the solution to actually producing actionable insights, which is a term I have heard now for Oh, I'm not going to tell you how many
years. You can guess. It's been a long
years. You can guess. It's been a long time. So, dives allow you to write
time. So, dives allow you to write dashboards to find out what's going on with your data pretty quickly.
What else is missing? Yes.
I was going to say that what the solution that we have is that when you when the schema changes, it creates a cookbook and that cookbook forces the um I'm going to have you say that on the
microphone.
Um and when a when the schema changes, it automatically creates a hook. That hook
forces our company to update all our MD files, all our markdown files that contain all the guides.
And the other thing that I'd say is missing part that we have added to this is another layer. Actually, my biggest frustration with context layer, semantic layer is that
there's it's like it's one layer. It's
actually always multiple layers. You
have two of them here. You should start off with a guide for your company. Like,
oh, a question comes in. Is it a live data question? Is it a look back
data question? Is it a look back historical question? What is the domain
historical question? What is the domain of this question?
Then the second layer which is, okay, now that I know that the domain or what's live and what's not live, then you get allocated to a domain and that domain has another guide to say, if it's
in this domain, then you're going to use this set of tables and you eventually get to the table. And the table um
table. And the table um context layer is a lot it's conflated with this semantic layer. Often times
they're one and the same.
So, you're talking about cascading levels of granularity.
Yeah, cuz your your you need your agent needs to know like what is this? According to what it is, what data do I need to look at?
Right.
And once I look at that data, how do I actually look at that data? How do I join them together? What are the different tables?
Yeah.
And then once I actually start to look at the where clause of my SQL statement, how do I know what the categories of a particular field mean? Which goes down to
field mean? Which goes down to the semantics.
YAML Yeah, yeah.
and semantic layer.
Okay.
Yes sir.
So, in an enterprise, they're sitting on top of all kinds of different sources, right? A lot of them are going to have
right? A lot of them are going to have legacy Hadoop files. Um
most of them will have a beta or still play or data bricks.
You've got a whole lot of different kinds of sources.
Each one of which is going to need a context layer. What happens
context layer. What happens when there is a discrepancy between the context layer or one source versus another? I mean, this is obviously
another? I mean, this is obviously helpful for for MotherDuck, but Mhm.
is it possible that someone's going to have to harmonize different context layers? That's the same thing as if
layers? That's the same thing as if you're harmonizing different data. It's
it's doesn't seem like it's all but it should It just pushes it away.
Uh I have a really simple solution for that.
Keep all your gold tier data in MotherDuck. Then you won't have that
MotherDuck. Then you won't have that problem.
I Yeah, this is obviously MotherDuck specific. There's no like big canopy
specific. There's no like big canopy over all of your data sources. That's
true.
And And that's a real problem.
People are are going to have to figure out where do you want your agents Where do you want to give people natural language access to agents? At what point
in your system? For me, I I would be fairly restrictive on what I would do. I
would keep that all in one place where I could control it.
I I'm you know, I've been Like I said, I've been around for a while. I have
been listening about AI my entire career right?
Um I have listened to the promises. I've
listened to the fears. I've listened to all of it. And what I'm watching now is yes, it's a phenomenal tool. I would
predict it. I would be in the last person in the world that become good friends with Claude.
But I see its usefulness, but I also see that you that as we move forward, there's going to be more and more um
desire to correlate to contain it, to make it do what we want it to do. So
your AI is working for you.
You're not working for AI.
It's a kind of different mindset and being a little more controlling.
Any other things that people have seen?
Yes sir.
Uh So like let's say like if you're asking a question about something like which is not in the guide, then it has to essentially go to the blog table at some point.
Hold on. I'm going to give this to you.
It's a good question.
So if you're asking a question about some metric that's not in the guide, uh so you have to essentially go down to the base tables. Where do the context of
base tables. Where do the context of that table come from? Like do you have to like write it manually or like You know, I I personally would actually
have a rule in my context layer that says if you don't know what I'm talking about, do not infer.
That you come back and say I it I was at This is just a little side story. I was at a meetup and uh the
story. I was at a meetup and uh the panel was talking about, you know, AI blah blah blah. And the last question was what does your What can't your agent
do right now that you wish it could? And
everybody was like saying a bunch of stuff. Do you know what my answer was?
stuff. Do you know what my answer was?
Say I don't know.
And I wanted to say I don't know. I
don't want it to guess. I don't want it to hallucinate. I don't want it it to
to hallucinate. I don't want it it to make me look bad if I [laughter] take its answer and give it to somebody. I
mean I mean, worst nightmare, right? So,
that's probably the one of the first rules I'd put in.
If you don't know what that term means, if you can't find something definitive, cuz you know when they're going through and creating the queries, they're telling you what they're doing.
And most of us like start it and go to another tab and just something else while it's working. I don't. I watch.
And I watch exactly who it's talking to, exactly where it's pulling data.
And then I'll go, "Oh, yeah. He has no clue. Look at all these sources he's
clue. Look at all these sources he's looking at and he can't find an answer."
So, it's say, "I don't know."
Then you know to go build your contact to to fix your context guide.
Anything else? Any other questions?
I tell mine to act like a woman cuz they will act they will ask for directions.
[laughter] That's great. You got to say that.
That's great. You got to say that.
All right. I said I tell mine to act like a woman because it will ask for directions.
Okay, any other questions? Any other
Oh, here we are.
Love your hat.
Thanks.
Do you guys have any tools to help with the context gang? Because like for example in the first presentation there was a whole like, "Oh, well, we can give the agent a hint like you should query
this way." And you have things where
this way." And you have things where it's like, "Oh, this this customer does this and we need to know this." But at a certain point you start bloating the context way down and you want to start
splitting things but not everyone knows how to do that and sometimes things like they slowly bloat. So, do you guys have any tools to help people figure out
like, "Okay, so this query will probably work 10 times better or be cheaper or be faster if you move it into a separate thing that it then has to go on and ask." Or "This thing is very too deep in
ask." Or "This thing is very too deep in the context layer. You should move it up the stack. Do you guys have any tools
the stack. Do you guys have any tools that would help recommend where something should be placed in the stack?
No, that's a really good point that you brought up and I'm glad that this is being recorded because our developers want to hear these kinds of of issues.
We want to hear what is your biggest frustration? What is What keeps you up
frustration? What is What keeps you up at night? If you're a data engineer and
at night? If you're a data engineer and you're telling your C-suite that yeah, they can ask you natural language questions. Do you hold your breath?
questions. Do you hold your breath?
You know, do you review everything before you hand it to them? Do you let them go into a board meeting, you know, [snorts] with AI-generated reports? Like, you
know, what are people doing? Oh,
Nathaniel's I'm tapping his watch, which means I have to let you go. So,
Nathaniel, you can wrap up. Oh, I think I have the thank you slide. Hold on.
He was waiting for my thank you slide.
Yay!
[music]
Loading video analysis...