Now livetag icon
The official Gen AI Landscape report for 2026 is here. Don’t miss out!Get your copybanner icon
sales insights iconSales Intelligence

The Data Engine Behind Rakuten Advertising’s Million-Dollar Pipeline

The Data Engine Behind Rakuten Advertising's Million-Dollar Pipeline

Sometimes we need a real-life example to know something is actually possible. Not a vendor demo, not a polished case study, but someone from your world showing their work, including the parts that didn’t go as planned. That kind of inspiration, proof that it can be done, is exactly what this webinar delivered.

Rakuten Advertising had a data problem most revenue teams will recognize: first-party data split across Salesforce and internal systems, third-party signals scattered across separate tools, and no reliable way to connect any of it into a prioritized, actionable queue. Prospecting a single account took over an hour. In June 2026, we hosted Julia Pan (Director of Demand Generation and Marketing Operations) and Aimee Pierce (Director of AI Adoption) from Rakuten, joined by our very own Jonathan Stefansky (VP, Sales Intelligence, Similarweb) on our Data Cooking Class webinar to walk through how they solved it. What they built generated $2.5 million in pipeline within six months, with no additional headcount.

The challenge: fragmented data makes every decision manual

Rakuten Advertising runs a performance marketing platform with a broad customer base. The outbound function depends on identifying the right companies, at the right moment, with the right message. On paper, they had plenty of data to do that. In practice, the data was everywhere and none of it talked to anything else.

Aimee Said in the webinar “Even though our bedrock is data, it’s disparate and it is everywhere Salesforce has this data and our database has this data, but they don’t necessarily talk to each other.” Sounds familiar?

We’ve been hearing this from clients a lot. This is the structural problem most RevOps and SalesOps teams recognize immediately: which data is trustworthy, which is actionable, and which is noise? Julia described the result as the “execution gap.” No system existed to translate raw, fragmented input into a prioritized, ready-to-work queue.

Studies show that poor data quality costs organizations money. An average of $12.9 million per year, according to IBM and Forrester research. A Validity survey of 602 CRM users found that 76% say less than half of their CRM data is accurate and complete. Rakuten wasn’t dealing with an edge case. They were dealing with the norm.

Layer one: the data foundation

The answer for them, had two layes.  Before any agentic workforce could function, the team needed a foundation of trustworthy, enriched account data. This is when they decided to  bring in Similarweb data.

“We used Similarweb to really understand, based on traffic, based on the sizing of that particular company, what prioritization did we want that to fall within the queue,” Julia said.

Web traffic data provided a behavioral signal on which target accounts were active and investing in digital. Company sizing signals helped segment the account universe by commercial weight, not just static firmographic proxies. Together, they gave the workflow something the CRM alone could not offer: an objective, current ranking of which accounts deserved attention and in what order.

Critically, Similarweb traffic data became the routing decision engine inside the workflow itself. “Similarweb is actually that deciding factor, what is the traffic, and that decides who gets that route,” said Aimee Pierce. Whether an account went into a fully automated outreach path or a human-reviewed path depended on the traffic threshold Similarweb’s data established.

The data foundation was not just input for salespeople. It was input for the agents. Without trusted external data at the routing layer, an agentic workforce collapses into a machine sending undifferentiated outreach.

To get there, Aimee shared a practical framework for any team evaluating which data sources belong in their workflow. It maps to the six questions we all learned for storytelling in school: who, what, where, when, why, and how.

  • Who owns this data, and who do you need in the room to understand it? Find the humans who own each system before you write a single line of logic. A 20-minute conversation with the right person can save weeks of rework.
  • What is the data built to do? Understand what the API can and cannot do, what edge cases exist, and what fields are actually reliable before you depend on them.
  • Where does it live, and where does it need to travel to be useful? Map every data source to its location. For Rakuten, that was CRM data in Salesforce and web signals in Similarweb, none of it in one place.
  • When was it last updated? Stale data does not produce mediocre results. It actively misleads the entire system at scale. Build time windows into your rules and flag anything outside them.
  • Why does this source belong in the workflow? Every data source you add is a maintenance commitment and a potential failure point. Rakuten added Similarweb because it answered questions the CRM could not. Everything else had to earn its place.
  • How is it flowing through the system, and how do you know if something goes wrong? This is what most teams fail to document. Julia built QA fields at every stage and metadata to trace everything. That is what makes data actionable, not just available.
Signal typePurpose in the workflow
Web traffic volumeAccount tier assignment and routing
Traffic trendsIdentifying growth-stage accounts
Company sizing signalsSegmenting outreach queues by commercial weight
Contact intelligenceEnriching CRM records with current data

 

The integration sat inside Salesforce, so data appeared in the existing workflow context without requiring a separate tool.

Layer two: the agentic workforce

With the data foundation in place, the team built an agentic workforce to automate and scale outbound execution.

“We basically, at a very high level, created an orchestration of agents to do outbound efforts on behalf of our very small BDR team,” said Julia Pan. The goal was to let agents handle research, enrichment, sequencing, and outreach while the team focused on review, judgment, and relationship-led follow-up.

But things are not as easy as one can think and Rakuten had to make mistakes in the way, to get to the perfect formula to success.

The V1 mistake: Rakuten’s first attempt was a single comprehensive agent handling the entire process. It didn’t work. A monolithic agent has to manage conflicting permissions across different Salesforce objects, hold too many context threads simultaneously, and fails in ways that are hard to diagnose. When it breaks, everything breaks. “It’s not just creating a contact. It’s not just ingesting that information. We needed to consider permission sets across all different objects, all different fields,” said Julia Pan.

The V2 shift: The revised design used specialized mini-agents, each responsible for one discrete task, each governed independently. One agent handles account research. One handles contact enrichment. One writes and sequences outreach. One logs CRM updates. Each has its own permission scope, its own failure boundary, and its own approval checkpoint. Errors become isolated and traceable, and the system scales by adding agent instances rather than rebuilding.

As Aimee Pierce put it: “It is literally peeling an onion. Every project we start I’m thinking, oh, this will be two weeks, and two weeks in I have discovered that no, it’s not the case at all.” Agentic infrastructure requires iteration. Build for learning, not for completeness.

The always-on context layer

What Rakuten built shows what’s possible when you get the data foundation and the agentic workforce right. But actually, there’s a third piece and it’s what separates a one-time build from a system that compounds over time.

The data foundation is not a set it and forget it setup step. It is a continuous layer of context

Market conditions change. Accounts shift. A company that ranked low in your scoring last quarter may be showing strong traffic growth and channel investment signals today. An ICP that made sense in January may need to be recalibrated by June. If your context layer is static, your routing decisions and your outreach decay in quality even as the workflow keeps running.

This is what Similarweb Sales Intelligence is built to address at the platform level. Three components work together to keep the context current:

Company Profile and ICP definition sets the foundation: who you are going after, what your offer is, and what signals matter for your specific sales motion. This is defined once centrally by RevOps, not left to individual reps to interpret.

Sales Signals and Scoring translate that definition into a living prioritization engine. Signals cover traffic changes, channel shifts, tech stack moves, leadership changes, funding rounds, and more. Scoring lets RevOps assign weight to the signals that matter for their ICP, so every account gets a ranked, context-rich priority score that reflects what is happening right now, not six months ago. Signal Management, launched in Q1 2026, lets admins set this centrally, so the definition of “high priority” is consistent across every rep, every tool, and every agent in the workflow.

Integrations are where the context meets execution. Via native CRM enrichment, the Signals API, and the Sales Intelligence MCP, that living context flows into Salesforce, into the agentic workforce, into the sequences, without anyone having to manually pull it. The context updates; the workflow responds.

As Jon framed it in the webinar, the change is continuous. You define your strategy once, embed it across your stack, and then revisit it as outcomes tell you what to adjust. The companies using this well are not building a workflow and walking away. They are building a system that gets sharper over time.

The results

Six months after deploying the data-enriched agentic workforce:

  • $2.5 million in pipeline created, with no additional headcount
  • Time to prospect a single account dropped from over an hour to a fraction of that, with research and enrichment handled upstream by agents
  • The outbound team shifted focus from administrative data work to review, approval, and relationship-led follow-up

Four takeaways for RevOps and SalesOps teams

1. Data quality before automation

You cannot automate your way past bad data. The agentic workforce only works because Similarweb’s external signals give the routing layer something trustworthy to act on. Cleanse the CRM first, decide which data matters, and make every record clean enough to hand off to the next agent in the chain.

Data quality before automation

2. Mini-agents beat monoliths

A one-size-fits-all agent does not work across complex GTM stacks. Define the smallest useful unit of work, build a dedicated agent for it, and add governance before adding scope. Each agent should have a clear responsibility boundary and an isolated failure mode.

3. Govern the stack

Decide which signals are worth acting on, then enforce those decisions consistently. Ungoverned workflows create CRM contamination and erode the data trust the whole system depends on. Set rules for what agents and humans can touch, and don’t retrofit governance after agents are already running.

4. Build a continuously evolving context engine

The work does not end at deployment. The context layer needs to be refreshed continuously: updating ICP definitions, adjusting scoring rules, and feeding new signals back into the workflow as market conditions change. Embed your strategy across your stack, and treat it as a living system, not a one-time build.

Build a continuously evolving context engine

What Rakuten built is replicable

The Rakuten story is not a one-off. Julia and Aimee weren’t operating with unlimited engineering resources or a purpose-built AI budget. They were solving a problem most revenue teams already have: too much fragmented data, too little time to act on it, and pressure to produce pipeline without adding headcount.

What made it work was sequencing. Trusted data first. Governance before scope. Mini-agents before monoliths. And a context layer that doesn’t sit still once the workflow is running.

The full session recording is available if you want to go deeper on any of the details. And if you’re working through any part of this yourself, whether that’s CRM enrichment, account prioritization, or thinking through how Similarweb data could fit into an agentic workflow, we’re happy to walk through it with you.

See which accounts are ready to buy

Prioritize your pipeline with real-time signals from Similarweb.

FAQ

What kind of data does Similarweb Sales Intelligence provide?

Similarweb Sales Intelligence provides external, behavioral data that first-party CRM data cannot replicate. This includes web traffic volume and trends, engagement metrics, channel mix, technographic signals, ad spend data, and company firmographics across over 100 million websites and 4 million apps in 190 countries. On top of that, Sales Signals surface timing-based alerts: traffic shifts, tech stack changes, leadership moves, funding rounds, new market entries, and intent spikes. The combination tells you not just who a company is, but what they are doing right now and whether this is the right moment to reach out.

How does Similarweb data get into our CRM and existing workflows?

There are several ways. Native CRM enrichment for Salesforce and HubSpot writes data directly into standard fields, so it is fully reportable and usable in existing workflows without iFrames or widgets. Enrichment jobs can be configured to run automatically on a daily, weekly, or monthly cadence with advanced logic (for example, enrich only new leads in a specific region). Beyond the CRM, data is accessible via the Sales Signals API, the Lead Enrichment API, the Contacts API, and the AI Outreach API. For teams building agentic workflows, the Sales Intelligence MCP connects Similarweb data to any LLM, AI agent, or automation tool in real time.

What is the Similarweb MCP and how does it work?

MCP stands for Model Context Protocol. Think of it as a universal plug that connects Similarweb data to AI tools in one setup rather than requiring separate API integrations for each. Once connected, any AI tool with MCP support, including Claude, ChatGPT, Cursor, and others, can pull live Similarweb data on demand. For RevOps teams building agentic workflows, this means agents can call Similarweb data as part of a multi-step task without manual exports or switching tools. Setup takes approximately five minutes.

Can RevOps teams define their own scoring and prioritization logic, or is it a black box?

RevOps owns the logic. AI Lead Scoring in Similarweb Sales Intelligence lets you define what “high priority” means for your ICP: which signals matter, what weight they carry, and what threshold separates high from medium from low. Every account gets a score and a “why now” brief that explains the reasoning, and scores sync directly into CRM fields. Signal Management, available to admins, lets RevOps set this centrally so the same definition applies across every rep, every tool, and every agent in the workflow. No rep builds their own scoring in a spreadsheet.

Can Similarweb data feed into our agentic workflows, or does it only work in the platform UI?

Yes, it works outside the platform. Via the APIs and MCP, Similarweb data can be called by any agent, automation tool, or LLM in your stack. In the Rakuten workflow described above, Similarweb traffic data was the routing decision engine: agents called it to determine which accounts entered automated sequences and which went to human review. The data is available as a live feed, not a static export, which means agents are always working with current signals rather than a snapshot from the last manual enrichment run.

 

Inbar Eytan photo

by Inbar Eytan

Product Marketing Manager

Inbar Eytan is the Product Marketing Manager for Sales Intelligence at Similarweb. With 8 years of experience in marketing, she believes in the power of turning data into actionable insights that make a real impact.

Related Topics
This post is subject to Similarweb legal notices and disclaimers.

Boost your consultative selling impact

Try Similarweb Sales Intelligence today — free of charge