欢迎光临
我们一直在努力

现任者即将到来 The Incumbents Are Coming The job is bigger than the record. __ A16Z

现有记录系统的优势在于,人工智能会使记录系统变得更加重要,而不是降低其重要性。为什么呢?因为理论上,客户现在可以将系统数据直接导入 Claude 或 Codex,而无需使用人工智能原生应用程序。

The Incumbents Are Coming – by Seema Amble – a16z

随着代理商通过这些系统完成更多工作,他们所掌控的数据和操作也变得更有价值。因此,现有的记录系统加上一个通用代理商(或由现有代理商开发的代理商)可能就足以完成工作。Salesforce 刚刚与 Anthropic 合作发布了Claudeforce ,它允许用户在Claude 内部使用 Salesforce CRM,而无需打开 Salesforce(至少他们是这么说的!)。

这种说法有一定道理。现任官员们并非无所事事。

现有企业有动机控制数据并收取数据访问费用,现在他们也更有动力推动自家代理的发展。我们已经看到,现有企业正在推广的产品不再仅仅是信息存储,而是能够承担更多实际工作。例如,DocuSign 的 Iris 用于审核合同,Atlassian 的 Rovo 旨在解决和路由请求,Klaviyo 的 Composer 用于构建营销活动并评估营销效果。这些产品促使现有企业不再仅仅满足于“简单地添加一个聊天机器人”,而是采取实际行动。

与此同时,Claude 开始协调跨应用程序的工作,使通用代理能够作为现有系统之上的一个层级发挥作用。Claudeforce 展示了这种模式的雏形:Claude 可以成为工作的入口,而 Salesforce 仍然控制着底层 CRM 数据和操作。界面和记录系统似乎正在解耦,这为人工智能原生创业公司围绕工作重新组合创造了机会。

那么,初创企业还能在哪些方面参与竞争呢?

这家垂直领域的AI原生公司必须通过专注取胜。它必须在执行特定的跨系统任务方面,比现有厂商的专用智能体或像Claude这样的通用智能体做得更好——后者能够利用底层工具重构任务。专注能够为这家垂直领域的初创公司带来更深入的访问权限、精心打造的数据资产、更具体的上下文信息,以及一个学习循环,从而随着时间的推移,更好地理解特定任务的优秀表现。

代理层级

要了解现有企业在哪些方面实力最强,以及垂直人工智能公司在哪些方面可能具有优势,思考四种类型的应用级代理会很有帮助,这些代理的区别在于它们各自需要多少自主性和判断力。

检索助手负责查找信息、总结文档、回答问题和撰写回复。流程代理执行基于规则的工作,例如更新记录或审批流程,他们会运用有限的判断力来理解任务并采取行动。策略代理将组织的操作手册、先例和阈值中的规则应用于模糊不清的案例。决策代理则对组织应该采取的行动做出判断,通常需要在战略、风险和资源之间进行权衡。

一年前,大多数现有企业的AI产品仍停留在检索阶段,甚至更糟。它们主要作为聊天机器人和分析工具,从现有记录中提取信息。如今,更多企业能够根据产品中已嵌入的工作流程、权限和规则采取行动。少数企业开始应用由人工定义的、范围较窄的策略。当决策依赖于现有企业管理记录之外的信息时,判断就变得更加困难。

我们看到的总体趋势是,现有企业正在向上攀升,从检索和处理转向通过政策进行更多判断,例如 DocuSign 的 Iris 系统就运用了法律团队的策略。但这种判断仍然很大程度上受限于现有企业过往的业绩记录。

这项工作比记录更重要

现有系统可以自动化处理更多围绕其核心记录的工作,但客户需要完成的任务远不止于此。这并非一道难以逾越的技术鸿沟,因为现有系统可以引入外部数据并添加工作流程,但他们的工作范围仍然会从他们已有的记录向外扩展

客户的工作是通过其人员、流程和软件的综合运用来完成的。合同并非法律文件,工单并非客户的解决方案,商机也并非销售。完整的工作涉及多个应用程序、团队,甚至公司。不同的参与方可能掌握不同的信息,并期望不同的结果。现有人员只能构建“本地工作系统”,但更大的机遇在于跨部门管理整个工作,包括所有参与方、人员、文档、决策、修订以及最终结果。

像 Claude 或 Codex 这样的通用代理或许能够访问所有这些系统,但仅仅拥有访问权限并不意味着它就能出色地完成任务。即使使用 MCP,跨应用程序提取信息也存在延迟,而且每个应用程序对同一客户、合同或交易的描述可能有所不同。即使通用代理可以访问一家公司内的所有应用程序,它也未必能自动访问外部各方(例如买卖双方)持有的信息。

即使存在这些限制,Claude 仍然可以凌驾于记录系统之上。Claudeforce 使 Claude 成为工作的入口,而 Salesforce 仍然拥有 CRM 数据和操作的控制权。此外,Claude 还可以通过与现有系统进行更深入的集成来不断获取上下文信息。

垂直人工智能如何学习

掌握更多工作细节至关重要,因为这能让初创公司了解最终结果背后的决策和调整过程。通过垂直整合的框架,结合合适的上下文、工具、工作流程和评估,模型在特定任务上的表现会更出色。学习循环正是系统(模型+框架)的精髓所在——每次任务完成后,系统都会利用智能体的工作、专家反馈和结果本身来改进下一次尝试。

仅凭完整的记录并不足以构成完整的培训课程。签署的合同表明了双方达成的协议,但并未列出所有考虑过的方案或所有例外情况的原因。已关闭的工单表明了问题的解决方案,但并未列出支持团队测试过的所有假设。

记忆与学习

实验室也在研发通用内存,因此仅仅保留上下文信息并不能带来持久的优势。记忆和学习是不同的。记忆可以帮助通用智能体回忆起客户的偏好或上周做出的决定。但它无法告诉智能体当时的工作做得好不好,专家为什么要修改,或者下次应该改进什么。而专注技术可以让公司看到更多同类工作的案例,更快地学习,并更好地衡量改进效果。

专业与机构

学习有两种类型。一种是学习优秀专业人士如何工作。公司可以通过专家设计的作业、真实的案例以及良好结果的标准来教授这种学习方法。

了解特定公司或客户的工作流程,包括其模板和先例、风险阈值和升级规则。这可以从客户现有的材料入手,并通过培训过程中的修正和例外情况加以改进。一些行业经验可以提升所有客户的服务质量,但针对某一家公司的经验可能只对该公司有效。

为了深入了解行业和机构,初创公司无需一开始就拥有海量的历史数据。它可以先针对行业制定一套课程体系,然后利用生产经验来了解机构运作。随着开放权重模型和训练后模型的增多,收集正确的数据和评估方法变得愈发重要。

长时间运行的智能体尤其受益于学习循环。当智能体执行数小时甚至数天的任务,经历包含众多中间决策和交接的复杂逻辑链时,学习循环的质量就显得尤为重要。这种长期任务可以分解,并通过一系列检查点进行评估。

需要明确的是,现有员工也会从他们各自产品内部的工作中学习。但这些学习仅限于他们负责的部分:例如 DocuSign 中的合同审核、Jira 中的工单解决或 Salesforce 中的销售管道更新。

哈维是如何炮制课程的

Harvey 在 Tenet 项目中的工作展现了垂直人工智能的发展路径。Harvey 利用合成数据、公开的法律数据和专家创建的数据,对一个开放权重模型进行了后训练,而没有一开始就使用客户数据。它创建了大约 1750 个法律任务环境,每个环境都模拟了合伙人分配的案件,包含完整的文档、工具以及用于确保高质量结果的专家评分标准。平均每个任务包含约 50 个具体标准。Harvey 没有等待数年积累足够的客户历史数据,而是构建了一套可用于训练和评估其模型的课程体系。

哈维更广泛的研究还表明,垂直行业公司如何训练模型来处理工作的不同部分,包括行业特定的能力,例如并购尽职调查和了解公司的全部知识。

这家初创公司的机会在于将所有要素整合起来。垂直领域的专业性有助于初创公司将工作分解为具体的功能,解决数据权限、系统和许可等方面的实际难题,并将产品融入客户的日常工作流程。最终,它就能坦然地承担起最终成果的责任。部署和责任感至关重要!

好的垂直人工智能市场需要具备哪些条件?

以 Harvey 为例,以下是评估潜在垂直人工智能市场的几个维度。

  • 专家能否快速判断人工智能的正确之处和错误之处,并解释如何改进它?

  • 这项工作是否难到需要运用判断力(而不是简单的规则)?

  • 这项工作是否经常进行,以便产品能够学习?

  • 初创公司能否从一个项目开始,逐步发展成为完成整个工作?

最好的垂直人工智能市场往往满足这四个条件,从而形成强大的学习循环。

法律、税务和会计领域的许多工作流程都符合这一标准,因为这些工作具有重复性,需要判断,而且结果已经过专家审核。这种模式也出现在一些不太明显的领域,例如工业领域。当出现制造缺陷时,质量工程师必须收集测试结果、供应商文件、设备日志和工厂规章,然后决定哪些是可接受的,哪些需要改进。这些信息可能分散在多个系统中。缺陷报告可能由某个现有系统存储,但没有哪个系统能够独立完成整个调查过程。初创公司可以从这项工作的一部分入手(例如撰写调查报告),并逐步掌握端到端的解决方案。为此,它需要学习优秀的质量工程师如何调查问题,以及特定工厂、客户或审核员如何做好这项工作。

这份工作目前仍然空缺。

有些工作只需要一套记录系统加上一个通用代理就足够了,但垂直领域原生人工智能初创公司仍然拥有巨大的发展机遇!最具潜力的垂直领域机会存在于那些工作频繁发生、专家判断至关重要且学习循环强大的领域。现有企业可能掌握着记录,实验室可能控制着入口,但垂直领域人工智能公司仍然可以通过自身在工作中做到最好而赢得市场。

The Incumbents Are Coming – by Seema Amble – a16z

— 

As agents do more work through these systems, the data and actions they control become more valuable. An incumbent system of record plus a general agent (or an agent built by the incumbent) therefore might be enough to get the work done. This is what Salesforce just announced in partnership with Anthropic: Claudeforce lets you work with the Salesforce CRM from inside Claude without having to open Salesforce (or so they claim!).

The argument is partly correct. The incumbents aren’t sitting idly.

The incumbents have an incentive to control and charge for access to their data, and now to push their own agents forward too. We’re already seeing incumbents market products that move from simply storing information toward doing more of the work itself. Docusign’s Iris reviews contracts, Atlassian’s Rovo aims to resolve and route requests, and Klaviyo’s Composer builds campaigns and assesses marketing performance. These products move incumbents beyond the “slap on a chatbot” strategy and into taking action.

At the same time, Claude is starting to coordinate work across applications, making it possible for general-purpose agents to act as a layer above the incumbents. Claudeforce shows what that can look like: Claude can become the front door to the job while Salesforce still controls the underlying CRM data and actions. The interface and the system of record look like they’re unbundling, creating an opening for AI-native startups to rebundle around the job.

So where can startups still compete?

The vertical AI-native company has to win through focus. It has to perform a specific cross-system job better than an incumbent’s purpose-built agent or a general-purpose agent like Claude, which will be able to reconstruct it from the underlying tools. Focus can earn the vertical startup deeper access, a deliberate data asset, more specific context, and a learning loop that builds a better understanding over time of what good work looks like for a given job.

The agent hierarchy

To see where incumbents are strongest and where vertical AI companies may have an advantage, it’s helpful to think about the four types of application-level agents, distinguished by how much autonomy and judgment each requires.

Retrieval assistants find information, summarize documents, answer questions, and draft responses. Process agents carry out rule-based work like updating records or routing approvals, using limited judgment to interpret the task, and act. Policy agents apply the rules in an organization’s playbooks, precedents, and thresholds to ambiguous cases. Principal agents make judgment calls about what the organization should do, often with open-ended tradeoffs around strategy, risk, and resources.

A year ago, most incumbents’ AI products were still at the retrieval stage, if that. They functioned mostly as chatbots and analytical tools pulling from existing records. Now, more can take action based on the workflows, permissions, and rules already embedded in their products. A few are beginning to apply narrow, human-defined policy. Judgment gets harder when the decision depends on information outside the record that the incumbent manages.

The broad shift we’re seeing is incumbents moving up the hierarchy, from retrieval and process toward more judgment through policy, like Docusign’s Iris applying a legal team’s playbook. But that judgment is still largely bounded by the record the incumbent owns.

The job is bigger than the record

An incumbent can automate more of the work around its core record, but the customer’s job to be done is bigger than the record. This isn’t a hard technical boundary because incumbents can pull in outside data and add workflows, but their work still expands outward from the record they already own.

A customer’s work gets done through a combination of its people, processes and software. The contract is not the legal matter, the ticket is not the customer’s resolution, and the opportunity is not the sale. The full job crosses applications, teams, and even companies. Different parties may hold different information, and want different outcomes. The incumbent is limited to building a “local system of work” but the larger opportunity is to manage the job across boundaries, including all of the parties, the people, the documents, decisions, revisions, and the final result.

A general purpose agent like Claude or Codex may be able to reach across all these systems, but access alone doesn’t mean it can perform the job well. Even using MCP, pulling info across applications has latency, and each application may describe the same customer, contract, or transaction differently. Even if a general purpose agent can access every app in one company, it doesn’t automatically have access to information held by external parties (e.g. between buyers and suppliers).

Even with these limitations, Claude can still sit above systems of record. Claudeforce makes Claude the front door to the job while Salesforce still owns the CRM data and actions. And Claude can keep gaining context through deeper integrations blessed by the incumbents.

How vertical AI learns

Owning more of the job matters because it lets startups see the decisions and corrections that produced the final result. A model performs better on a specific job via a vertical harness with the right context, tools, workflows and evals. The learning loop is what improves the system – the model + harness – after each job, using the agent’s work, expert feedback, and outcomes themselves to improve the next attempt.

Completed records alone don’t automatically suffice as a training curriculum for the work. A signed contract shows what the parties agreed to, but not every alternative considered or every reason for an exception. A closed ticket shows the resolution, but not every hypothesis the support team tested.

Remembering vs. learning

The labs are also building general-purpose memory, so preserving context alone is not a durable advantage. Remembering and learning are different. Memory can help a general purpose agent recall a customer’s preferences or a decision made last week. But it doesn’t tell the agent whether the work was good or why an expert changed it or what to change next time. Focus lets the company see more examples of the same job, learn faster, and get better at measuring improvement.

Profession and Institution

There are two kinds of learning. Learning how a good professional does the work. A company can begin teaching that with assignments created by experts, realistic examples, and standards for a good result.

Learning how a particular firm or customer does the work, including its templates and precedents, risk thresholds and escalation rules. That can begin with the customer’s existing materials and improve through corrections and exceptions in the training process. Some lessons about the profession can improve the product for every customer, but lessons about one firm may just improve the product only for that firm.

To learn both the profession and the institution, the startup doesn’t need to start with the largest stock of historical data. It can manufacture a curriculum for the profession, then use production to learn the institution. As open-weight models and post-training increase, assembling the right data and evals becomes even more important.

Long running agents especially benefit from learning loops. When an agent does work over hours or days, through a complex logic chain with many intermediate decisions and handoffs, the quality of the learning loop matters even more. The longer horizon task can be unbundled and evaluated through a set of checkpoints.

To be clear, incumbents will also learn from the work performed inside their products. But that learning will cover only their part of the job: the contract review inside Docusign, the ticket resolution inside Jira, or the pipeline update inside Salesforce.

How Harvey manufactured the curriculum

Harvey’s work on Tenet shows what the path for vertical AI can look like. Harvey post-trained an open-weight model using synthetic data, public legal data, and human-expert created data, without using customer data to start. It created roughly 1,750 legal-task environments, each simulating a partner-assigned matter, complete with documents and tools and an expert rubric for a high-quality result. The average assignment had about 50 specific criteria. Instead of waiting years to accumulate enough customer history, Harvey manufactured a curriculum it could use to train and evaluate its model.

Harvey’s broader work also shows how a vertical company can train a model for different parts of the job, including industry-specific capabilities like M&A diligence and understanding a firm’s total knowledge.

The startup’s opportunity is to put it all together. Vertical specificity helps the startup to break the work into specific capabilities, tackle the practical frictions around data rights, systems, and permissioning, and fits the product into the customer’s daily work. Eventually, it can feel comfortable taking responsibility for the result. Deployment and responsibility are key here!

What makes a good vertical AI market?

Using Harvey as one example, here are a few axes for evaluating potential vertical AI markets.

  • Can an expert quickly tell what the AI got right or wrong, and explain how to improve it?

  • Is the work hard enough that judgment matters (vs. simple rules)?

  • Does the work happen often enough for the product to learn?

  • Can the startup begin with one assignment and grow into doing an entire job?

The best vertical AI markets tend to meet all four, giving them a strong learning loop.

Many workflows in legal, tax, and accounting meet this test because the work repeats, requires judgment, and experts already review the results. This pattern also shows up in less obvious markets like industrial work. When a manufacturing defect appears, a quality engineer has to gather test results, supplier documents, equipment logs, and the plant rules, then decide what’s acceptable and what needs to change. That information may sit across several systems. One incumbent may store the defect report, but no single product owns the full investigation. A startup could begin with one piece of this work (e.g. drafting the investigation report) and over time own the end-to-end resolution. To do this, it would need to learn both how a good quality engineer investigates a problem and how a particular plant or customer or auditor does it well.

The job is still up for grabs

There are jobs where a system of record plus a general-purpose agent will be enough, but there’s still meaningful opportunity for vertical AI-native startups! The strongest vertical opportunities exist where the work happens often, expert judgment matters, and the learning loop is strong. The incumbent may own the record and the lab may control the front door, but the vertical AI company can still win by becoming the best at doing the job itself.

赞(0)
未经允许不得转载:171主机测评 » 现任者即将到来 The Incumbents Are Coming The job is bigger than the record. __ A16Z
分享到: 更多 (0)

评论 抢沙发

  • 昵称 (必填)
  • 邮箱 (必填)
  • 网址