All insights
Systems

Skills vs System: A Head-to-Head on One Enterprise Fintech

Systems9 min read

Anthropic account-research skill vs Sales Kriya agent fleet on one live deal

Anthropic's account-research skill will set your enterprise sales team up to burn its domain.

Not because the skill is bad. Because it works confidently, on stale data, and hands your reps a list of names to email at a target company where those names do not work anymore.

We ran it head-to-head against our own agent fleet last week on a real live-pipeline account. Same input. Same target. A public, Vista-backed enterprise fintech we are actively pursuing.

The skill recommended cold-emailing a Chief Revenue Officer whose appointment was announced 32 months ago and a Chief Strategy Officer who is not in the current employee roster. If we had trusted the output, the emails would have bounced, our domain reputation would have taken the hit, and the buyer we already have three touchpoints with would have watched us look sloppy.

Here is what happened on both runs.

What the packaged skill did

Data sources it actually reached: web search only. The Apollo people-search integration returned an API_INACCESSIBLE error on the free plan. The HubSpot company lookup returned zero records. Gmail, Drive, and Calendar were not part of the skill's execution flow, so they were never queried.

That is the entire sourcing story. Web pages plus two failed integrations.

Working from that, the skill produced a company profile, a "buying centre map" with six named people, a qualification read, and a recommended entry point.

Then it noted, verbatim, at the bottom of the output: "No prior relationship history was retrieved, because no relationship source was queried."

What our agent fleet did

Our SCOUT and WEB agents ran on the same account, in the same run window. The stack:

  • Live LinkedIn roster verified today via Apify. Every named stakeholder is either confirmed against today's LinkedIn or flagged as "verify before contacting."
  • CRM record for this account, including the deal history, past objections, and touchpoint timestamps.
  • Email and call history across the last quarter.
  • ACCOUNT_MEMORY from prior SCOUT, WEB, EDGE, and PLAYBOOK runs on the same account.
  • Web research as a supplement, not the primary source.

Then twenty-plus agents on one shared deal memory produced a buying-centre map, a champion-building plan, a next-action recommendation calibrated to the specific silence pattern in the deal, and a coverage-gap flag.

The differences

1. The skill recommended cold-emailing a Chief Revenue Officer whose appointment was announced in December 2023

The packaged skill listed a specific person as Chief Revenue Officer, Economic Buyer, medium confidence. Sourced from a BusinessWire press release dated December 2023. That release is almost three years old.

Our WEB agent running against the live LinkedIn roster today does not find that person in the current employee list. He may have left. He may still be there. The skill has no way to know, because it does not reach the live roster.

The skill flagged this in a "known gaps" section: "Whether the CRO is still Chief Revenue Officer." That is honest. It is also not useful when the recommendation built on that unverified name is: reach out to him.

2. The skill recommended a Chief Strategy Officer as the entry point

Second-choice entry recommendation from the skill: a Chief Strategy Officer and General Manager India Operations.

Our WEB agent flagged the same person in its "not confirmed in current roster, verify before outreach" section. That flag was auto-generated from the Apify roster comparison against prior memory. The person was in a previous WEB run. The person is not in today's roster.

Cold outreach to a name that has left the company equals "you have the wrong person," which equals burned credibility with the buying centre you actually need to reach. A skill without a live-roster feed cannot catch this.

3. The skill missed the actual champion candidates

Six confirmed current stakeholders our WEB agent surfaced that the skill did not:

  • VP Enterprise Sales, based in Bengaluru, 2 years in role, no engagement history, no baggage from previous sales conversations. Primary champion candidate.
  • SVP Sales, Mid Market, 3 years in role, five years at the company, aware of the internal build's limitations from the inside.
  • AVP Strategic Initiatives, Office of CEO, promoted three weeks ago. This person is the CEO's operating proxy on strategic initiatives. Any project the CEO is personally tracking, including the pilot we already ran, will likely route through her.
  • VP Digital Transformation, four months into a mandate that maps exactly to internal capability building for AI operations.
  • Chief Technology Officer, internally promoted eight months ago. The technical gate on any integration decision.
  • VP Cloud, Platform, and Security Engineering, three months into the VP role, running a thorough security review because his credibility depends on it. New-VP-in-prove-himself-mode.

The packaged skill named none of these people. It named a marketing chief who is not in the current roster and a strategy chief who is also not in the current roster.

4. The skill invented a CMO

The skill listed a specific person as Chief Marketing Officer, medium confidence, as an influence-layer contact.

Our WEB agent, running against today's roster, lists three separate marketing leaders at the company: an SVP Marketing (internally promoted 11 months ago), and two VPs of Marketing. None of them share the name the skill produced. The name the skill produced does not appear anywhere in the current employee list.

That is not a small failure. That is a hallucination on a named stakeholder in a buying-centre recommendation. If a rep runs the skill and emails that person, the email bounces or gets forwarded internally with "who is this vendor and why are they writing to a person who does not work here." One data point, one burned bridge.

5. The skill had no CRM context

Verbatim from the skill's output: "No prior relationship history was retrieved, because no relationship source was queried."

Our SCOUT had all of it. A 60-minute discovery call with the founder on the fifth of July. An in-person demo with the SVP Sales the next day. A WhatsApp follow-up on the tenth, no reply. Verbatim objections quoted from the calls, including the founder's "not scalable" and the SVP's "good for small companies." A read on which objections are sunk-cost defences of an internal build and which are actual product rejections.

The skill treated this account as if it were day one of research. It is not. This deal has three touchpoints and forty-five days of silence.

6. The skill's rep-facing recommendation set us back, not forward

Read the skill's opening hook out loud: "You changed your commercial model in February. Outcome based pricing means your sellers now have to negotiate success criteria instead of a licence, and your historical win rate data no longer predicts anything."

That is a diagnosis of their sales organisation, delivered to the person the skill correctly identified as sensitive to that framing. The skill even flagged the anti-pattern in the same document: "Never lead with diagnosis of their sales organisation." Then it led with a diagnosis of their sales organisation.

Our SCOUT's opening hook, informed by the actual call history, targeted a specific stakeholder with a calibrated question ("When you say 'not scalable,' what does scalability look like for a 1,200-rep org?") and a specific proof point ("we have a comparable-scale customer case study").

The skill wrote a generic pitch. The system wrote the next move in a live deal.

Why this is a category difference

A skill queries the web through a language model whose training data is already stale on the day you use it.

Claude Fable 5's knowledge cutoff is January 2026. Any leadership change, funding event, product launch, or organisational reshuffle that happened at your target company after that, the model does not know from its training. It backfills with a web search, but the web itself lags. Google indexed the December 2023 press release announcing the CRO's appointment. It may not have indexed his exit, if there was one. LinkedIn changes to job titles do not propagate to public search for weeks. Corporate about-us pages are updated on a quarterly cadence at best.

Put together, on the day we ran this, that gave us:

  • A Chief Revenue Officer sourced from a 32-month-old press release
  • A Chief Strategy Officer as entry point who is not in the current employee list
  • A Chief Marketing Officer name that does not match the current roster
  • No knowledge that the deal already had three touchpoints and forty-five days of silence
  • A pitch hook that violated the anti-pattern the skill itself named

A system queries the deal, not the web. That means the verified live roster (pulled from LinkedIn today, not the model's training window), the CRM history, the email and calendar traces, the timestamped call objections, and the prior agent runs. Every stakeholder is either confirmed against today's LinkedIn or flagged as "verify before contacting."

Both cost tokens. The skill is free. The system is not.

For a solo operator running a light research task on a company they have never talked to, the skill beats doing nothing. It surfaces the outcome-based-pricing signal from February, it reads Tracxn and Inc42 correctly on headcount and funding, and it gives you a starting frame.

For an enterprise sales team about to send an email that arrives at a CRO who left the company, a strategy chief who left the company, and a CMO whose name may never have been at the company, the skill costs more than the tokens. It costs the deal.

The consistency test

There is a second thing to name. Anthropic's packaged skill runs on the caller's Claude session. Session context bleeds. Prior deals, prior prompts, prior context all shape the output.

Ask five reps on your team to run the same skill on the same account on their own Claude instances. Same prompt.

You will get five different buying-centre maps.

That is not a Claude problem. That is a skill-versus-system problem. A skill is a session. A system is a fleet of agents on one shared deal memory. When the rep changes teams, the deal memory stays. When the rep quits, the deal memory stays. When five reps ask the same question, they get the same answer.

At one person, session-scoped is fine. At a team of ten, it is chaos. At a team of a hundred, it is board-level forecast risk.

What this means if you are evaluating

If you are running a small team, use Anthropic's account-research skill for the first pass on cold accounts. It is free, it works, it beats doing nothing.

If you are running an enterprise sales function with a live pipeline, do not let free skills near the accounts that matter. The failure modes stack fast: stale data on senior titles, missing stakeholders in the current roster, invented names in marketing and executive rows, no CRM context, session-scoped output that varies rep to rep.

At the price of a burned first email to the wrong person on your top account, the free skill is not free.

What we built instead

Sales Kriya runs twenty-plus agents on one shared deal memory. The agents write to and read from the same account row. Every stakeholder is verified against a live LinkedIn roster before it is surfaced to the rep. Every touchpoint is timestamped. Every objection is quoted verbatim from the call transcript.

That is the system. It is not a smarter prompt. It is a stack that keeps the deal memory alive across reps, across weeks, and across every agent that touches the account.

If your team is running Claude skills in production and you want the head-to-head on one of your live accounts, we will do the same comparison run for you. Reach out through the /apply page and we will schedule the diagnostic.

Anthropic's packaged skill referenced in this comparison is sales:account-research on Claude Academy, launched on 20 August 2026. The comparison run was executed on 24 August 2026 against a public enterprise fintech we are actively pursuing. The two Sales Kriya outputs (SCOUT and WEB) were produced by our production agent fleet in the same run window.

See whether the model fits your team

We embed a senior Sales Leader and an AI System of Action to build, enforce, and run your playbook across 6 to 12 months. Start with a revenue diagnostic.

Get a revenue diagnostic

Discussing this on LinkedIn.

We open every blog up for discussion on the Sales Kriya LinkedIn page. Share your take there, or send us a private reaction below.


Or send us a private reaction.

We read every one. If something here rang true (or not) for your team, tell us. No newsletter, no sequence.

We never add you to a list. One human reads this.

More from Sales Kriya

Sales Kriya is an embedded sales coaching partnership plus an AI System of Action for enterprise B2B teams. The full model is on how we work, and the four-condition fit check is on the alignment matrix. If a stalled deal of your own came to mind while you were reading, the three free tools. Score a deal, Score a call, Pick a play. Are the closest thing to seeing the model run on your data.

Keep reading