推荐微信扫码登录以获得更好的体验!另:🎉建议更新到最新版,以体验新功能!🎉 版本更新记录→

【AI前沿】CaaS崛起:Agentic AI的上下文即服务

字幕摘录

时间英文中文
0:13So, hi, everyone.大家好,你们好。
0:14Thank you so much for taking the time to join this session.非常感谢你们抽出时间参加本届会议。
0:16I hope I'll, or at least I can guarantee you, I'll do whatever it takes to make it worth我希望我会,或至少我可以保证, 我会不惜一切代价使它变得值得
0:20your time.你的时间。
0:22My name is Omel.我叫奥梅尔
0:23I lead the product marketing team over at Bright Data.我在Bright Data公司领导产品营销小组
0:26Just by maybe a quick show of hands.也许只是举手
0:27Who here is familiar with Bright Data?谁在这里熟悉亮数据?
0:31Okay.摆
0:31We can do better.我们可以做得更好。
展开字幕全文(386 条)
序号英文中文
1So, hi, everyone.大家好,你们好。
2Thank you so much for taking the time to join this session.非常感谢你们抽出时间参加本届会议。
3I hope I'll, or at least I can guarantee you, I'll do whatever it takes to make it worth我希望我会,或至少我可以保证, 我会不惜一切代价使它变得值得
4your time.你的时间。
5My name is Omel.我叫奥梅尔
6I lead the product marketing team over at Bright Data.我在Bright Data公司领导产品营销小组
7Just by maybe a quick show of hands.也许只是举手
8Who here is familiar with Bright Data?谁在这里熟悉亮数据?
9Okay.摆
10We can do better.我们可以做得更好。
11I'll pass it on to our brand team.我会把它传给我们的品牌团队。
12Bright Data是一家网络数据公司.
13Basically we help more than 20,000 teams around the world, including more than 70% of the基本上我们帮助了全球两万多支球队,包括超过70%的
14world's biggest AI labs, to extract data from the web.世界上最大的人工智能实验室 从网上提取数据
15Just to put this in perspective of what scale we're talking about, we're talking well over仅仅从我们所说的规模的角度来说,我们谈得很好
1650 billion pages, HTMLs every day, more than 20 petabytes of video, audio, and other media500亿页, 每天 HTML, 超过 20 petabytes 的视频,音频,和其他媒体
17data.数据。
18So that's just the perspective of the type of work that we do at Bright Data.这就是我们在Bright Data公司的工作类型。
19But enough about us.但是关于我们已经够了。
20Personally, I joined Bright Data about three years ago, which essentially gave me front就我个人而言,我三年前加入了 Bright Data, 这基本上让我站在前面
21row seats at everything around AI and the web and how they started to actually connect.在AI和网络周围的所有位置排成一排, 以及它们是如何开始连接的。
22It sounds very old, but if you think about it, only maybe less than two years ago, right,听起来很老了,但是如果你考虑一下, 也许不到两年前,对吗,
23we were able to start using, access the web, and search the web through cloud or through我们得以开始使用、访问网络并通过云或通过网络搜索网络
24闲聊GPT.
25That option didn't even exist in the earlier versions, right?这个选项在早期版本中甚至不存在,对吧?
26So that connection, that way in which both AI and the web are starting to converge is因此,这种连接, AI和网络开始融合的方式是
27something that is still evolving and evolving rapidly.一些仍在迅速演变中的东西。
28That's part of what I want to try and shed some light on today and talk about a new emerging这也是我想尝试的一部分 并给今天一些启示 并谈论一个新的新兴
29breed of companies that's coming out of this connection.由这种联系产生的公司
30I think we can basically agree that the web is by far the world's greatest source of data.我认为我们基本上可以同意, 网络是世界上最大的数据来源。
31At least historically, when it comes to Bright Data, that's all we cared about, right, helping至少在历史上,当它涉及到亮数据, 这就是我们关心的,正确的,帮助
32our customers extract data from the web.我们的客户从网上提取数据
33But with the emergence of AI and the emergence more recently of AI agents that need to do但随着AI的出现 和最近出现的AI代理 需要做
34knowledge work, right, the web is no longer just a source of data.知识工作,对了,网络不再仅仅是数据的来源.
35We can actually start looking at it as a source of context.我们实际上可以把它看成是背景的来源。
36Context in the sense that if I do knowledge work and I have knowledge agents that support也就是说,如果我做知识工作,我拥有支持的知识代理人
37my work, I want to go out to the web, find the information that I need, use it as context,我的工作,我想去网络, 找到我需要的信息,用它作为背景,
38but keep on working.但继续工作
39So the data itself is only a step in the process for something bigger, for the actions I need因此,数据本身只是进程的一个步骤 对于更大的东西, 对于我需要的行动
40to take, for the conclusions I need to draw, and for every downstream application that对于我需要得出的结论,对于每一个下游应用
41follows.接下来。
42The first one is to figure this one out, oh, sorry, even before that.第一个是想出这个, 哦,对不起,甚至在那之前。
43But one thing that I want all of us to bear in mind because this is going to follow us但有件事我希望大家记住 因为这会跟着我们
44through the rest of this conversation.通过其余的谈话。
45The web is messy, it's unstructured, and most importantly, it changes all the time.网络杂乱无章,结构不整齐,最重要的是,经常变化.
46This is a chart that shows data decay, right, that it's analysis done by our team that shows这是一张显示数据衰减的图表,没错, 它是我们团队所做的分析显示
47data decay basically how long after a new page, a new piece of content goes live, it数据衰减 基本上在一个新页面之后多久 一个新的内容开始运行,它
48is no longer relevant, right?已经无关紧要了吧?
49So social media, it's easy for us to understand, it's far less than a day.所以社交媒体,我们很容易理解,还不到一天.
50But also news, finance, retail, 30 days later, data that was collected, it's mostly no longer但新闻,金融,零售,30天后, 收集到的数据, 大部分已经不再
51relevant.相关。
52And when we acknowledge that, this simple notion, we understand that extracting context而当我们承认,这个简单的概念, 我们理解提取的背景
53from the web or relying on the web is not a snapshot, it's not a one-time effort, it's从网络或依赖网络不是快照,不是一次性努力,而是
54not even a monthly effort.连月费都没有
55It's something that we need to keep on doing.这是我们需要继续做的事情。
56It's something we need to look at as an ongoing process and something that we need to be mindful我们需要将它视为一个持续的过程 和我们需要注意的东西
57The first ones to figure it out were, of course, search companies, right?最早发现的当然是搜索公司,对吧?
58Only three years ago, we were on the far left, right?仅仅三年前,我们在最左边,对不对?
59This is in our lifetime, right?这是在我们一生中,对不对?
60Three years ago, we were on the far left, everything was Google.三年前,我们在远方, 一切都是谷歌。
61There was complete and total dominance up until three years ago for the past 20 or so过去20年左右的三年前 完全和完全的主导地位
62years.岁月
63That's what we're talking about.师便打曰.
64Purely for humans, search something, go, collect the information you need, and carry on.纯粹为人类,搜索一些东西,去,收集你需要的信息,然后继续.
65Then, fast forward, maybe one, half years ago, two years ago, search began to appear然后,快速前进,也许一年半前,两年前, 搜索开始出现
66within the LLMs, within the chatbots, right?在LLMs内部, 在聊天室内部,对不对?
67Which already started blurring the line between humans and agents because now the same bots已经模糊了人类和特工之间的界限 因为现在都一样
68also have the same web search available, so the same LLMs have the same access through也可以使用同样的网页搜索,所以同样的LLMs通过
69API for the bots, so we started seeing that convergence happening, right?蛋蛋的API, 所以我们开始看到 交汇的发生,对不对?
70And so, for the first time, we're seeing more and more traffic flowing down, search traffic,所以,我们第一次看到 越来越多的流量流下,搜索流量,
71search intent flowing down these channels, not only to Google.搜索意图流出这些频道, 不仅仅是谷歌。
72And last but not least, right, we now see a whole breed of companies, the AI search最后但并非最不重要的是,没错,我们现在看到的是 整个品种的公司,AI搜索
73companies, which you're familiar with them.公司,你熟悉它们。
74I caught a talk yesterday by Will, the CEO of Exxon, and Parallel, and U.com, and Tavili,我昨天接到了埃克森公司总裁威尔 和平行公司和U.com公司 以及塔维利的谈话
75and a bunch of others.还有一群人
76They are purely built in indexing the web, especially for agents, not even looking at它们纯粹是用来编制网络索引的 特别是针对代理商 连看都不看
77the humans involved anymore.与人类有关
78So Google's dominance when it comes to search, if Google was synonymous of web search, that所以谷歌在搜索时的主导地位 如果谷歌是网络搜索的同义词的话
79is very much shaking.非常颤抖。
80When there's blood in the water, the sharks come.当水中有血,鲨鱼就来了.
81Just last week, Amazon, Amazon announced, I don't know how many of you saw it, that就在上周,亚马逊,亚马逊宣布, 我不知道有多少人看到它,
82they developed their own index and started allowing the ability to retrieve data from他们开发了自己的索引,并开始允许从中检索数据的能力。
83the web, to retrieve context for agents on agent core.网络,以检索代理核心上的代理上下文。
84Amazon developed their own search engine.亚马逊开发了自己的搜索引擎.
85Two weeks before that, it was Microsoft.两周前是微软公司
86Microsoft always had skin in the game.微软在游戏中总是有皮肤.
87I don't know, one or two percent of the world's search traffic went to Microsoft.我不知道,全球搜索流量的一二成都投给了微软.
88But they have now repackaged it and launched it, again, as part of Web IQ, as part of their但他们现在重新包装 并再次推出它 作为Web IQ的一部分
89suite for agentic development and orchestration.用于代理开发和管弦乐的套房。
90So we're seeing more and more this space becoming crowded.我们看到这个空间越来越拥挤。
91But when we're talking about context, and we're looking at this through the lens of但当我们谈论背景时, 我们从镜头看这个
92search, I believe it only tells us part of the story, right?搜索,我认为它只告诉我们 故事的一部分,对不对?
93I can search for, I don't know, what's the cost of a certain pair of sneakers this morning,我可以搜索,我不知道, 什么是成本 某双运动鞋今天上午,
94right, on a certain website.对,在某个网站上
95I cannot really search for how has that price changed over the last six months, what discount我实在找不到过去六个月价格的变动
96it had, right?曾经是 对吧?
97我能找到Bright Data的公开职位
98We do.没错
99I urge you to go have a look.我劝你去看看
100But I can't see how that was a chart and how that changed over time and how the headcount但我看不出这是一张图表 和它如何变化 随着时间和人口统计
101of the company changed over time.公司随时间而变化。
102And all of that information existed on the web simply back then.所有的信息都存在于网络上 就在那时。
103So when we actually start to think about it, we understand that there's much more context所以,当我们真正开始思考的时候, 我们明白,还有更多的背景
104in the web than what web search allows us to extract.在网络中比什么网络搜索让我们可以提取。
105And this is what we started seeing in the recent years, a whole new breed of companies这是我们近年来开始看到的 全新的公司
106rising.上升。
107We like to call them internally CAS, context as a service, because that's what they do.我们喜欢在内部称之为CAS, 上下文作为一种服务, 因为他们这样做。
108They allow agents to tap into them, MCP, CLI, just pure good old API, and actually start他们允许特工挖掘他们,MCP,CLI, 只是纯粹好的 API,实际上开始
109extracting data to retrieve data so they can reason over for whatever knowledge work they提取数据以检索数据,以便他们能够为任何知识工作提出理由
110are responsible for.负责。
111We see this happening in e-commerce.我们在电子商务中看到这种情况。
112We see this happening in travel.我们在旅行中看到这种情况。
113We see this happening in finance, in market research, in HR, in real estate, in a bunch我们从金融、市场研究、人力资源、房地产、一堆
114of other domains.其它领域。
115I'll show a few examples in a second, right?我马上会举几个例子,对吗?
116What all of these have in common is that they don't only just discover the web, you know,这些东西的共同点是 它们不只是发现网络
117in terms of think crawling, think searching, think all of that, accessing, extracting the思考爬行,思考搜索, 思考这一切, 访问,提取
118data and indexing it.数据和索引。
119They take it a step further.他们更进一步。
120They actually develop knowledge graphs to start structuring all of the entities and他们实际上开发了知识图 开始构建所有实体和
121to dedupe them.驱除他们。
122And they start enriching them with a lot of different sources.他们开始用许多不同的来源来丰富它们。
123So they actually start merging all of that data.所以他们开始合并所有的数据。
124If you think about it, they kind of behave like vertical search engines, right?想想看,他们的行为就像垂直搜索引擎,对吧?
125They are a very, very, very good search engine for something very specific.它们是一个非常,非常,非常好的搜索引擎 对于非常具体的东西。
126And it's already in full motion, right?并且它已经完全运动了,对不对?
127So as I said, we see this in finance, in market research, in retail, in e-commerce, in GTM,因此,正如我所说,我们看到这一点在金融,市场研究,零售,电子商务,GTM,
128in sales intelligence.在销售情报。
129What all of these companies, by the way, have in common, they're all part of Bright Data's顺便说一句,所有这些公司的共同点 都是Bright Data公司的一部分
130startup program if you are a builder.如果您是构建者, 则启动程序 。
131And this is a hot space to go in because I think we're only tapping the surface.这是一个热的空间 进入,因为我认为我们只是 敲打表面。
132I invite you to scan this and apply up to $20,000 in credits and all sorts of co-marketing.我邀请你扫描一下 并申请最高2万元的信用和各种共同营销。
133But that's enough self-promotion.但是,这已经足够自我促进了.
134So CAS as an industry is already in full bloom.所以CAS作为一个行业已经完全开花。
135And as always with these situations, right, also the traditional players aren't left too和往常一样 传统球员也不剩了
136much behind.相当落后
137There at the bottom, you see good old data as a service, you see ZoomInfo, right?在底部,你看到好的老数据 作为一种服务,你看到ZomoInfo,对不对?
138By researching for this presentation today, I also saw that they launched that thing at通过研究今天的演讲,我还看到,他们推出的东西在
139the top.顶部。
140其名称为GTM.AI.
141You can only imagine how much they paid for that domain.你只能想象他们为此付出了多少钱
142But they launched a secondary brand for ZoomInfo that is catering specifically for the need但是他们推出了ZoomInfo的二级品牌 专门满足需求
143of agents.特工人员。
144Look at the wording, right?看看措辞,对不对?
145They talk about GTM work, right, that knowledge work, that research that you do when you need他们谈论GTM的工作,对,知识的工作, 研究,当你需要的时候
146to prospect, when you need to do headhunting, whatever it is you need to do that involves当你需要猎头的时候 不管你需要做什么
147people mostly, straight from cloud code, straight from codecs or any other agent.大部分人,直接从云码, 直接从解码器或任何其他剂。
148They understand the gap, right?他们理解差距,对不对?
149So yeah, it's fine to think of CAS as an evolution of DAS, and it is in a way, but it's catering所以,这是很好的认为 CAS 是DAS的进化, 它在某种程度上,但它的餐饮
150与Google不同。
151When agents need them, it's different than people.当特工需要他们时,它和人不同.
152When we let this one sink, that at the very least we have two different types of paths当我们让这个沉下去的时候 至少我们有两种不同的路径
153to complete knowledge work as an agent, we can start thinking about this in terms of作为代理完成知识工作,我们可以开始思考这个问题
154web context engineering.网络上下文工程.
155We can start thinking about this in terms of how do I optimize for the specific task,我们可以从我如何优化具体任务的角度来思考这个问题,
156more importantly when things come as it is, how do I optimize this for breeds of tasks?更重要的是,当事情到来,它是什么, 我如何优化这个品种的任务?
157How do I do this for various parts of the organization that I'm building for?我如何为我建立的组织的各个部分做这个?
158If I'm an AI engineer, I need to serve different teams, they may have different needs.如果我是AI工程师 我需要为不同的团队服务 他们可能有不同的需要
159It's very tempting to throw AI search at all of them, but maybe that's not optimal.将AI搜索扔到他们身上是很诱人的,但也许这不是最佳的.
160Maybe I need a combination of both.也许我需要两者的组合
161Maybe I can start seeing all sorts of cost efficiencies emerge from that.也许我可以开始看到 各种成本效率的出现。
162So for the second half of this presentation, we actually went ahead and created a test.因此,对于这个介绍的后半部分,我们实际上做了一个测试。
163This is not a benchmark.这不是一个基准。
164You won't see any something concrete that I can say with great confidence other than你将看不到任何具体的东西 我可以非常自信地说
165the actual research that we did because we wanted to start unraveling the different considerations我们所做的实际研究 因为我们想要 开始解开不同的考虑
166and how do these two stack up against each other.这俩人怎么互相对抗
167So we designed a test.所以我们设计了一个测试。
168We went for something basic.我们去找一些基本的东西。
169We said, okay, let's take a company, an entity, and try and enrich it across 25 different我们说,好吧,让我们采取一个公司,一个实体, 并尝试并丰富它 超过25个不同的
170fields.字段。
171Some of them are very easy, you know, the company domain, the name, the headquarters,其中一些非常简单,你知道,公司域名,名称,总部,
172but some are more challenging, right, things about hiring and people and something.但有些更具有挑战性,对, 有关雇用和人的东西。
173And we built a simple agent, a loop that uses Opus 4.8 as the harness, and it starts我们建造了一个简单的代理,一个循环 使用Opus 4.8作为绳索,它开始
174to go field by field, go out, search for it, or retrieve it from the CAS, do it again and逐个出场,出去搜索,或者从CAS中取回,再做一次
175again and again until it completes and brings back, sets some guardrails, you know, like一次又一次 直到它完成和带回来, 设置一些护栏,你知道,喜欢
176budget and stuff just to keep it fair.预算什么的 只是为了保持公平
177And I'm going to share with you the results.与你分享结果
178So the first thing that we would care about, right, being knowledge work would be, sorry,所以,第一件事,我们会关心, 对,作为知识工作是,对不起,
179we ran it 100 times on all of the sponsors of today's event.今天所有赞助商上百次
180So the first thing that we saw in terms of coverage is that there's pretty good convergence.因此,我们首先看到的是, 在覆盖面方面, 存在着相当好的趋同。
181They all did fairly well, right?他们都做得很好,对不对?
182I'll get to the two at the bottom in a second.我马上到底部两个
183So search were consistent performance, one of the major CAS providers were also very所以搜索是一致的, 一个主要的CAS提供商 也非常
184well.不错
185The third one, by the way, you can see on Locker and SERP, SERP is good old data, good第三个,顺便说一句,你可以看到 在洛克和SERP,SERP是好的旧数据,好的
186old Google.旧谷歌.
187We basically did the same thing just with Google, and it performed pretty well in extracting我们基本上做了同样的事情 与谷歌, 它的表现相当不错 提取
188that information.那个信息
189Native is Claude's own search, and you see that they converge really well.原生地是克劳德自己的搜索, 你可以看到,它们真的汇合得很好。
190I was originally surprised about the two CAS solutions at the bottom.我最初对下面的两个CAS解决方案感到惊讶.
191It was counterintuitive.这是反直觉的。
192I expected CAS to dominate this thing because that's, you know, you had one job, right,我期望CAS能主导这件事 因为你有一份工作
193to map out these companies.以规划这些公司。
194But after diving into it a bit more, you understand that, well, they are limited in the sense但潜水多一点后,你就会明白,从意义上来说,它们有限
195that they know what they have about an entity.他们知道自己对于一个实体有什么感觉。
196If I ask them the question that is beyond that, they will never have that data, right?如果我问他们一个超出这个范围的问题, 他们永远不会有这些数据,对不对?
197Unlike a search, it can go out and continue searching and exploring it.与搜索不同,它可以外出继续搜索和探索.
198If they didn't collect data about the recent job hiring, it will never be there, right?如果他们不收集有关最近招聘工作的数据,那就永远也不会存在,对吧?
199So it makes sense that they are a bit behind, but I'm sure at the same time that they have所以说他们有点落后 但我确定他们同时
200a lot of other advantages that we simply didn't ask for, a lot of other fields that they didn't我们根本没要求的很多其他优势, 很多其他的领域,他们没有
201have that aren't represented.没有代表。
202So again, it creates some complexities on how do we measure coverage when it relates因此,它再次造成一些复杂因素,说明我们如何衡量涉及
203to the specific job that we need to do rather than in general.以完成我们所需要的具体工作,而不是一般工作。
204The second thing we looked at was cost, of course.我们看到的第二件事当然是代价。
205Here we started seeing it spread out a bit.在这里,我们开始看到它扩散了一点。
206You can see that massive bulk in the center.你可以看到中央那块大块
207Most of the search and the CAS and even using Google, right, converged to pretty much the大部分搜索和CAS,甚至使用Google,对, 汇合到几乎大部分
208same cost, only different, right?同样的成本,只是不同,对不对?
209The CAS was just about the service itself, what you pay the vendor, right?CAS只是关于服务本身, 你付给供应商什么,对不对?
210All of the other search solutions, you also needed a lot of token burn to actually structure所有其它的搜索解决方案,你还需要很多符号燃烧来实际结构
211that data so you can actually act on it and use it as something retrievable, right?数据,这样你就可以 实际采取行动,并用它 作为可检索的东西,对不对?
212So it's the same output.故同输出.
213Native, obscenely expensive, and the CAS on the right, I'm sure you're all familiar with,本地人 淫秽昂贵 右边的CAS 我肯定你们都熟悉
214the barf are the most expensive in the industry.烤肉是这个行业最贵的
215I will not name and shame them.我不会给他们起名和羞辱
216Interesting, you see that small CAS there at the left, that CAS number two, they were有趣的是,你看左边那个小CAS, 第二CAS,他们是
217very cheap.非常廉价。
218And they're also the ones that are here at the bottom, which is funny because what I他们也是最底层的人 这很有趣 因为我
219believe is happening there is that we're seeing, even within this industry, niche players that我们正看到 即使在这个行业里 也有合适的角色
220have lower quality data but much cheaper, they're already carving that niche of the数据质量较低,但价格便宜得多 他们已经在刻画了
221long tail, right, of small shops or small usage so that they don't want to pay as much长尾巴,右,小商店或小用途 以免他们想支付那么多
222and don't need as much data.不需要那么多数据
223And we're all seeing them branch out there.我们都看到他们的树枝在那里。
224Most of you here, I presume, are engineers.我想你们大多数是工程师
225So there's a very evident question that we did not ask here, which is, what is the one所以有一个非常明显的问题,我们没有在这里问, 那就是,什么是
226thing that an engineer would care about?工程师会关心的?
227Thank you.谢谢
228Let's talk about scale.我们来谈谈规模
229This is the cost, not for the whole hundred, this is the cost per one, for one record.这是成本,而不是整个100, 这是每个成本, 记录。
230What happens if we need a million?如果我们需要一百万怎么办?
231Now yes, a million records will not fit in the context we're obviously, we're not talking现在,是的,一百万的唱片 将不符合的背景 我们显然,我们不是谈论
232about a single run that needs a million.单程需要一百万
233You can think about a million in terms of the frequency, right?你可以考虑一百万的频率,对不对?
234If I am a market research, I do the diligence for private equity, I revisit these companies如果我是市场研究,我做私募股权的尽职调查,我再看看这些公司
235all the time.所有的时间。
236I ask more questions about them as the time goes by.随着时间的流逝,我问更多关于他们的问题。
237Was there any new news about them?他们有什么新消息吗?
238Was there anything that changed?有什么变化吗?
239Did somebody join?有人加入吗?
240Did somebody leave?有人走了吗?
241Do they have new hires?他们有新工作吗?
242I keep on asking the same thing.我一直问同样的事情。
243So when I'm talking about this, the multiply by a million, it's not just about the number所以当我谈论这个,乘以一百万, 它不仅仅是数字
244of companies, it's the frequency in which I'm asking it.公司,这是频率 我问它。
245Frequency is the cost killer when we talk about these, and we need to acknowledge that,频率是成本杀手 当我们谈论这些, 我们需要承认,
246right?对吧?
247We're thinking about this in terms of web context engineering.我们在网络环境工程方面考虑这个问题。
248We're starting to look at it differently.我们开始不同看待它。
249Every repeated query costs the same as the first.每次重复查询的费用与第一次相同.
250Even if it brought back the exact same answers, nothing changed, pay up, right?即使它带来了完全相同的答案, 没有什么改变,支付,对不对?
251No, it's false positives, for sure go in.不,这是假阳性, 当然进去。
252Token costs, right?托肯成本,对不对?
253We saw in the model that there's a very high token.我们在模型中看到 有一个很高的标志。
254We know that doesn't shrink well over time.我们知道这不会随着时间而减弱
255There's always some volume element in terms of the cost, but it's not the same as flatlining,成本方面总有一些量元素,但与平铺不同,
256right?对吧?
257And if we bring this back to knowledge work, this is where we see teams that are starting如果我们把这个带回 知识工作, 这就是我们看到的团队开始
258to cut corners.切开角落。
259So I won't research this company every day, I'll look at it once a week or once a month.所以我不会每天研究这个公司,我会每周或每月看一次.
260I won't ask that question now.我现在不会问这个问题
261I don't want all the results, I'll only take 10 results, 20 results, something.我不想得到所有的结果,我只拿10个结果,20个结果,什么的.
262So we already have the setup.所以我们已经安排好了
263We have what we need to do the knowledge work, but at the same time, we're not extracting我们有需要做的知识工作, 但与此同时,我们没有提取
264all of the value because we're starting to be conscious about cost, right?所有的价值 因为我们开始意识到成本,对不对?
265Basically we're renting context.基本上,我们租了背景。
266We're not owning the context that we use.我们不拥有我们使用的背景。
267That is a very important distinction.这是一个非常重要的区别。
268Again, if we're good engineers and we ask ourselves what about scale, the second most再说一遍,如果我们是优秀的工程师, 我们问自己,什么是规模,第二
269obvious thing that will come to mind now, so how about we build it?显而易见的事情,现在会想到, 那么我们如何建造它?
270What if we take all of that web data ourselves and stick it in some vector database and try如果我们自己把所有的网络数据 粘在某个矢量数据库里 然后尝试
271and see what comes out of it?看看里面有什么
272So I asked my engineer to do exactly that.所以我要求我的工程师这样做。
273Again, this is a test.再说一遍,这是一个考验。
274This is not a benchmark or a full-blown operation.这不是一个基准或全面的行动。
275This is a day's work at best just to illustrate the concept and to show something about the这是一个一天的工作,充其量只是 说明这个概念 并展示一些关于
276cost efficiencies that you can generate potentially by doing it yourself, potentially, in specific成本效率,你可以 潜在的通过自己做到这一点, 潜在的,具体
277scenarios.假设
278The test, simple.测试,很简单。
279Take the company name, nothing but run it through Google, find the relevant entries,取公司名称,只是通过谷歌运行, 找到相关的条目,
280the relevant URLs of that company in various websites that have all of that information.拥有所有信息的各种网站上的公司的相关URL.
281You use search when you don't know the source, but when we're talking about company enrichment,你用搜索 当你不知道来源, 但当我们谈论公司浓缩,
282we all know these sources.我们都知道这些来源。
283We all know where that data comes from.我们都知道这些数据来自何处。
284The cost, also bring it from them, zoom in for bringing it from them.成本,也从他们带来, 放大从他们带来。
285It's the same thing over and over again.亦复如是. 师曰.
286Why not just go straight to the source?为什么不直接去源头?
287Why are we doing that middleman thing?我们为什么要做中间人的事?
288LinkedIn公司,LinkedIn 工作,Crunchbase, 在那里,我们有刮刮机的人
289just tap in and you start paying as a pay-as-you-go.随便你便开始付钱
290We built two dedicated scrapers.我们造了两台专用刮刀
291我们有一个新的AI工具,叫Scraper Studio.
292It basically lets you build a scraper for any website in less than five minutes, all powered它基本上让你为任何网站 建造一个刮刀 在不到5分钟,所有电源
293by AI, and then it also has a self-healing function.由AI进行,然后它也具有自愈功能.
294If the website changes, it fixes itself and keeps on going.如果网站有变化,它会自我修复,并继续运行.
295Merge it all into one entity, basic heuristics.把它们合并成一个实体 基本热力学
296If there's conflict, choose that over that.如果有冲突, 选择这一点。
297Eventually, we have a data set of these 100 companies, zero AI cost involved.最终,我们有这100家公司的数据集,零AI成本。
298There's no tokens.无有征兆.
299Coverage, fairly well.覆盖,相当好。
300Not amazing, not the best that we saw here, but stacking up pretty well.并不令人惊奇,也不是我们所看到的最好的, 但堆积得很好。
301Again, this is just a day's experiment, probably not even as much.再说一遍,这只是一天的实验, 可能甚至没有那么多。
302Again, very specific tasks, very limited context, very limited situation.同样,非常具体的任务、非常有限的背景、非常有限的情况。
303Tread lightly and then proceed with caution when it comes to conclusions.轻轻地绊倒,然后在得出结论时谨慎行事.
304The real story is not this.事实并非如此。
305The real story is this.真实的故事就是这个。
306That's what it costs to just go and fetch that data that is out there.这就是去获取外面的数据的代价。
307We think about knowledge graphs.我们考虑的是知识图表。
308We think about entities, but if you think about, for example, LinkedIn,我们想的是实体,但是如果你想, 例如,LinkedIn,
309the data is already structured in form of entities.数据已按实体形式排列。
310There's an entity for a company, there's an entity for a person,公司有实体,人有实体,
311there's an entity for a job, and they're connected between them.有个实体来工作 他们之间有联系
312Sometimes the ontology is already there.有时本体论已经存在.
313Again, this is not the most complicated of scenarios, but this is pretty damn good.再说一遍,这不是最复杂的情景, 但这是相当他妈的好。
314Now, yes, it took time to set up.现在,是的,它需要时间 设置。
315So it's not really fair to compare apples to apples when it comes to the cost,所以,如果把苹果和苹果比起来 成本是不公平的,
316because these are out of the box.因为这些都出柜了
317You can just tap into the API.你可以进入API。
318That one that I just showed you required some setup.我刚刚给你看的那张 需要一些设置。
319Let's say it's a week.如是说礼拜.
320Let's price it at $5,000 just to give us some perspective.让我们用5000美元的价格来给出一些视角。
321We can actually start thinking about this in terms of a tipping point.我们可以从一个临界点开始思考这个问题。
322We can actually start thinking about what is that tipping point我们可以开始思考什么是临界点
323in which it makes more sense for me to build it myself than keep on renting it.我个人建造它比继续租更合理
324Now again, everything to the left of that dot, in this case,再说一遍,这个点左边的一切,在这种情况下,
325it was just over 15,000 entities or queries when we think about it.我们考虑的时候,只有15,000多个实体或查询。
326So it made sense to do it at this point.因此,现在这样做是有道理的。
327Maybe it's not 15.也许不是15岁
328Maybe it's 30.也许是30岁
329Maybe it's 100,000.或谓十万.
330Maybe it's 10,000.也许是一万块
331It really depends on the use case.这真的取决于用例。
332But there is a tipping point in which it actually makes sense to do it yourself,但有一个临界点, 它实际上有道理自己做,
333which leads us to the fact that both AI Search and CAS and all these solutions,这导致我们发现 AI搜索和CAS 以及所有这些解决方案,
334they're very good in the sense that you can just plug and play.他们很好,因为你可以 插和演奏。
335But if your knowledge work needs are persistent and consistent,但如果你的知识工作需要 持续和一致,
336and to a certain degree may even continue escalating and growing,并在某种程度上可能继续升级和增长,
337then this is perhaps a direction to start considering.那么这也许是开始考虑的方向。
338Maybe I can just go ahead and build my own.也许我可以自己建
339Because the nice thing about it is that all of the things that we see on the left因为我们左边看到的一切
340up until the third part is upfront investment.直到第三部分是前期投资。
341And the most important thing that whatever retrieval happens later on from the agents最重要的事情是,不管后来从特工那里收回什么
342is free.是免费的。
343Not really free, but you get what I mean.不是免费的 但你懂我的意思
344There's no added cost.没有额外的成本。
345I can just ask that question over and over again.我可以反复地问这个问题。
346I did not like the first answer.我不喜欢第一个答案。
347I'll ask it again.又问.
348I'll ask it 100 times until I get what I need.我会问100次 直到我得到我需要的。
349I have no more fear, no more cutting corners, which is maybe the most important thing.我不再害怕 不再割角 这也许是最重要的
350And I'm leaving aside the fact that this is also custom business logic.我撇开这个事实, 这也是定制商业逻辑。
351I can connect it with my own data.我可以用我自己的数据连接它
352There's all sorts of other advantages of owning it.拥有它还有其他各种好处.
353We'll keep it to the imagination.我们会保持它的想象力。
354Remember, we asked about a million, not about 15,000.记住,我们问了一百万,不是15,000。
355This compounds.这种化合物。
356This compounds greatly.这非常复杂。
357We need that horizon.我们需要那个视野。
358Remember, the web keeps changing.记住,网络一直在变
359We saw the staleness of the data and how the data decays.我们看到数据的停滞 和数据如何衰减。
360So we need to be thinking about this in the long run and how this will evolve所以我们需要从长远的角度来考虑 这个问题会如何演变
361when we keep on asking the questions about the entities that we care about.当我们不断问我们关心的实体的问题时
362Just to wrap it up.只是把它包起来。
363So AI, Search, CAS, they can get you very far when what you need is ad hoc所以,AI,搜索,CAS, 他们可以让你非常远 当你需要什么 临时
364and what you need is always changing.你需要的总是在改变
365When sometimes you look at different things,有时你看着不同的东西,
366even the mix and match of them for certain tasks use this,即使混合和匹配 某些任务使用这个,
367for certain tasks use that.用于某些任务。
368You can, I'm sure, again, I just saw the test.你可以,我敢肯定,再次, 我刚看到测试。
369There's a lot of ways to optimize it just like any other context engineering有很多方法可以优化它 就像其他环境工程一样
370and use lighter models and use other stuff.并使用更轻的模型 并使用其他的东西。
371There's a lot of great stuff to be done.众生种种妙法.
372But eventually, the frequency will come and bite you in the ass但最终,频率会来咬你的屁股
373when it comes to cost.当涉及到成本。
374And that's something to be mindful of.这是值得注意的。
375And there's a fair chance that that tipping point is much lower than you think.并且有相当的机会, 这个临界点 比你想的要低得多。
376And that's something that, as we design these systems,当我们设计这些系统时,
377when we think about web context engineering, we need to be mindful of that.当我们考虑网络环境工程时,我们需要注意这一点。
378And last but not least, the last slide we showed.最后但并非最不重要的是,我们展示的最后一张幻灯片。
379Owned context compounds while rented decays.租赁衰变时拥有上下文化合物。
380It's not a one-time task.非一时之任.
381Again, if it's a one-time question, use AI, Search, it will be amazing.复次若是一时问,用AI,Search,会令人惊叹.
382When you need to do it over and over again,当你需要一次又一次地做的时候
383there's a fair chance that it will not,很有可能不会
384that you're missing out on potential compounding effect你错过了潜在的复合效应
385and you are losing out.你输了
386Thank you very much.谢谢
该视频共有字幕 386 条。解锁更多字幕为会员功能,请移动到 价格