Skip to content
Multi-Channel Outcome Dataset · Available for License
Built to GDPR standards · Privacy safe · No PII

Ground truth for AI agents that sell.

Not synthetic. Not a contact list. Real outreach outcomes mapped into a privacy safe schema.

15M+ outbound touches. 5 years. LinkedIn + email. Mapped to sender, message, CTA, reply, and outcome.

01
0
Verified Booked Calls
02
0
Years of Continuous Collection
03
0M+
Outbound Touches Analyzed
04
0K+
Human Replies Captured
05
0+
Structured Data Fields
Live · Daily Ingest
+10,000 emails / day+1,000 LinkedIn touches / day~3.65M new labeled touches added every year
Every licensee snapshot includes the latest ingest through the delivery date. Choose daily, monthly, or quarterly refresh.
Five years. One source of truth.
Interactive Sample Explorer

Explore Privacy-Safe GTM Training Data

Select a use case and see the type of structured sample output SmartReach can provide for AI agents, GTM engineering, revenue intelligence, and human-in-the-loop workflows.

Privacy Safe Derived Data Only Sampled From 13.5K Booked Calls · 2K Email Replies · 4.5K LinkedIn Actions · 1K Buyer Matches No Names, Emails, LinkedIn URLs, Company Names, or Raw Messages Shared
Verified Meeting Outcomes

13,504 booked calls analyzed.

Every booked call is tagged by industry, geography, company size, and decision-maker title, so your AI learns from outcomes that match your ICP.

What This Dataset Captures

The full outreach interaction · not the contact.

SmartReach AI captures who was targeted, what type of sender contacted them, what channel was used, how the message was structured, whether the CTA was hard or soft, whether a reply was received, what outcome occurred, and what action should happen next.

Each record is privacy safe and structured around lead category, sender profile category, channel, subject-line structure, message structure, CTA type, reply status, normalized outcome tag, scoring, and next-best-action labels.

Lead category
Sender profile
Channel
Subject structure
Message structure
CTA type
Reply status
Outcome tag
Next best action
Why This Is Different

Most GTM datasets tell you who a prospect is. This one shows what happened after outreach.

The sender profile layer is a key differentiator · buyers can study how sender credibility and profile signals move real campaign outcomes.

01

Real Campaign Outcomes

Records connect lead category, sender profile, message hook, CTA type, channel, reply status, and final outcome. AI agents can learn from what actually happened · not synthetic examples or raw contact records.

02

The Sender Profile Layer

Study how sender credibility, profile strength, region, follower band, connection band, age band, and gender or race/ethnicity category affect replies, scheduling, booked meetings, unsubscribes, and negative outcomes.

03

Multi-Channel Outcome Tags

LinkedIn workflow status and email response events unified under one normalized outcome vocabulary · booked, in talks, scheduling, not interested, unsubscribed, bounce, auto-reply, needs review, and more.

04

Not Contact Data

No emails, phone numbers, LinkedIn URLs, company names, or raw private reply text. Category-level fields and anonymized sender IDs · evaluate and license without exposing raw contact records.

The Bottom Line

Lead + sender + channel + hook + subject + message + CTA + outcome + next best action.

The full interaction · normalized into one schema across LinkedIn and email.

Dataset Layers

Six layers per record.

Every interaction is mapped into a shared six-layer schema where applicable.

L1

Lead Intelligence

Seniority, role category, region, ICP segment, industry category, company size band, and buyer-role classification.

L2

Sender Intelligence

Sender profile category, age band, gender category, race or ethnicity category, region, follower band, connection band, authority score, and profile strength.

L3

Email Infrastructure

Sender type, email domain type, domain extension, ESP category, mailbox age band, domain reputation band, warmup status, and sending volume band.

L4

Message Intelligence

Message hook, message length, tone, opening structure, question count, personalization level, CTA structure, value proposition type, and reading time.

L5

Subject-Line Intelligence

Subject length, theme, structure, question flag, personalization flag, tone, specificity score, and reply-style signal.

L6

Outcome Intelligence

Reply received, booked, scheduling, in talks, qualified, future interest, not interested, not relevant, unsubscribed, bounce, auto-reply, no-show, needs review, and next best action.

Full Schema

Explore the Email and LinkedIn data dictionaries.

Two privacy safe schemas. Field-by-field definitions, types, and examples. Toggle between datasets.

Ideal Licensee

Built for teams shipping AI that sells.

Four buyer profiles. One dataset. Each one uses a different slice of the same 93+ field schema.

AI SDR Platforms

Train reply classifiers and next-message policies

Use 110K+ real human replies mapped to intent, sentiment, objection, and outcome. Fine-tune models that route, respond, and escalate the way top human SDRs do.

GTM AI Startups

Fine-tune next best action models

Sender profile, message, CTA, reply, and booked call outcome in a single row. The exact schema needed for reinforcement learning on outbound decisions.

Revenue Intelligence Vendors

Benchmark outbound performance across industries

13.5K verified meeting outcomes across 5 years, segmented by industry, seniority, geography, and channel. A defensible baseline for your customers' dashboards.

LLM Labs

Labeled B2B communication data at scale

Real outbound copy, real replies, real outcomes. Clean labels for supervised fine tuning, evaluation sets, and safety testing on business messaging.

Vs. The Alternatives

Why nothing else gets you here.

The four paths a team usually considers before licensing. And what breaks on each one.

Scrape LinkedIn
  • ·Violates ToS and GDPR
  • ·No outcome labels
  • ·No sender or CTA context
Apollo, ZoomInfo, Clay
  • ·Contact records, not outcomes
  • ·No reply text or intent labels
  • ·Nothing to train a policy on
Synthesize with GPT
  • ·No ground truth
  • ·Model collapse on your own outputs
  • ·Cannot evaluate against reality
In-house collection
  • ·12 to 24 months to reach scale
  • ·Single-industry bias
  • ·No historical baseline
SmartReach Dataset
The only path that actually trains a selling agent.
  • First-party, consent-based collection
  • 15M+ touches mapped to real outcomes
  • 93+ labeled fields per row, both sides of the conversation
  • 5 year historical baseline across industries
From Reply To Training Row

One real reply. Twelve labeled fields.

This is how a raw human response becomes a training row your model can actually learn from.

01 · Raw Reply
"Thanks for reaching out. This is interesting but our budget cycle resets in January so nothing new is getting signed off this quarter. Feel free to circle back mid Q1 and we can look at a small pilot then."
Anonymized. Real production reply.
02 · Derived Labels
reply_intentinterested_needs_timing
reply_sentimentpositive
objection_typebudget_cycle
buying_stageproblem_aware
seniority_confirmeddecision_maker
next_best_actionoffer_q1_pilot
recommended_cadence_days45
escalate_to_humanfalse
outcome_labelmeeting_booked_delayed
confidence_score0.86
channelemail
reply_stagereply_2
03 · Training Row
{
  "input": {
    "sender_role": "founder",
    "prospect_seniority": "VP",
    "message_variant": "intro_softCTA_v3",
    "reply_text_tokens": [...]
  },
  "labels": {
    "intent": "interested_needs_timing",
    "objection": "budget_cycle",
    "next_action": "offer_q1_pilot",
    "outcome": "meeting_booked_delayed"
  },
  "reward": 0.72
}
Ready for SFT, DPO, or eval sets.
Last refresh:November 2026·Next:February 2027
Collected from first-party outbound campaigns run by SmartReach AI and consenting partners. No scraping. No purchased contact lists.
What Buyers Can Learn

One shared taxonomy. Comparable answers.

Buyers can learn which lead types, sender profiles, channels, message hooks, subject-line structures, CTA types, and follow-up paths produce replies, scheduling, booked meetings, unsubscribes, bounces, not-interested responses, and not-relevant outcomes.

Compare LinkedIn and email under one outcome taxonomy · and test whether sender profile data, message structure, CTA type, and channel improve outcome prediction and next-best-action decisions.

LinkedIn Learnings

Which lead types respond, which sender profiles perform, which hooks work, and which workflow states should trigger follow-up, routing, or suppression.

01 / 05

Email Learnings

Which subject lines, sender types, domain types, ESPs, message structures, and CTAs produce replies, unsubscribes, bounces, or meeting intent.

02 / 05

Sender Learnings

Which sender profile categories produce replies, scheduling, booked meetings, not-relevant responses, or negative outcomes.

03 / 05

Cross-Channel Learnings

How LinkedIn and email compare by reply status, outcome tag, CTA type, message hook, sender type, and lead segment.

04 / 05

AI Decisioning Learnings

When an AI agent should follow up, wait, send a booking link, route to sales, suppress, mark not relevant, or request human review.

05 / 05
Built For

If you're building AI that sells across LinkedIn and email, this is your training data.

01

AI SDR Platforms

Train agents that pick the right sender, channel, hook, subject, and CTA · grounded in real multi-channel outcomes.

02

Sales Automation Tools

Ground your sequencing and next-best-action logic in normalized outcome tags across LinkedIn and email.

03

Outbound AI Agents

Learn the sender effect · how credibility, authority score, and profile band change reply and booking rates.

04

AI Model Evaluation

Test models against pre-derived message and subject-line features, including hook, tone, CTA structure, personalization level, and specificity.

05

Conversational AI

Classify replies: bounce, auto-reply, human, positive intent, clarification, negative · labeled at record level.

06

Revenue Intelligence

Enrich pipeline scoring with fit score, AI-agent fit, confidence, and next-best-action labels per record.

Shared Outcome Vocabulary

One normalized outcome tag across every channel.

LinkedIn replies and email replies get labeled with the same vocabulary · so a model can score, route, and evaluate them side by side.

dataset.smartreach / outcome_tags
LINKEDIN · EMAIL · UNIFIED
Normalized outcome tags
01Booked
Positive
02In talks
Positive
03Scheduling
Positive
04Connection accepted
Neutral positive
05Qualified
Neutral positive
06Future interest
Neutral positive
07Sent link
Neutral
08Objection handle
Neutral
09Clarification question
Neutral
10Pending connection
Pending
11Awaiting reply
Pending
12Not interested
Negative
13Not relevant
Negative
14Unsubscribed
Negative
15No show
Negative
16Bounce
System failure
17Auto-reply
System neutral
18Excluded or failed
System
19Needs review
Review
6 layers · 1 row
L1Lead

Region, seniority, department, ICP segment, industry, company size band, redacted role category.

L2Sender profile

Anonymous sender key, authority score, follower + connection band, region, age band, replies + bookings produced.

L3Channel + workflow

LinkedIn or email · pending, accepted, awaiting reply, replied, failed, excluded · last action time.

L4Message + subject AI features

Hook type, opening structure, tone, personalization level, CTA structure, subject theme, question flag, specificity.

L5Normalized outcome

Shared outcome tag across channels · outcome category · sales stage · outcome source.

L6Scoring + next best action

Fit score, AI-agent fit, confidence, data-quality flag, next best action label.

Why this matters

One vocabulary means LinkedIn and email are directly comparable · for scoring, routing, evaluation, and next-best-action modeling.

refresh · quarterlyoutcome categories · positive · neutral · negative · pending · system failure · system neutral · reviewschema · v2026.Q2
Refresh Cadence

Choose your refresh cadence.

New data is added every day. 10,000 email touches and 1,000 LinkedIn touches, labeled and mapped to the same schema. Choose how often that new data ships to you.

Each refresh gives buyers fresh examples of what is working now across channels, sender profiles, message structures, CTA types, reply patterns, unsubscribes, bounces, and booked-meeting outcomes.

Cadence
Daily
~11,000 new touches / day
For teams retraining agents continuously.
Cadence
Monthly
~330,000 new touches / month
Standard for production AI SDR stacks.
Cadence
Quarterly
~990,000 new touches + full re-label
Included in the base annual license.
What every refresh includes

New LinkedIn and email outcome records

Freshly processed outreach interactions mapped to the shared schema.

Updated outcome labels

Booked, scheduling, in talks, not interested, not relevant, unsubscribed, bounce, auto-reply, needs review, and other normalized tags.

Sender-performance refresh

Updated sender profile summaries, authority scores, profile strength bands, and outcome performance by sender type.

Message and subject-line updates

New message hooks, CTA structures, subject-line patterns, tone labels, personalization fields, and performance signals.

Fresh scoring fields

Updated fit scores, agent fit, confidence scores, data quality flags, and next-best-action labels.

Evaluation Package

Paid evaluation dataset.

A structured package to test ingestion, outcome classification, sender-performance signals, message and subject-line features, next-best-action logic, and model lift · before committing to the full license.

The evaluation package is designed for internal testing, schema review, model benchmarking, outcome classification, and next-best-action analysis. Production use, resale, redistribution, and public release require a separate annual license.

Request Evaluation Dataset
What's included
  • 01CSV and JSONL formats
  • 02Shared outcome taxonomy
  • 03Data dictionaries
  • 04Sender profile summary
  • 05Message and subject-line AI feature guide
  • 06Field coverage report
  • 07Evaluation terms
Annual License

The full dataset, for a year.

The standard annual license provides access to the larger privacy safe LinkedIn and email outcome dataset · with quarterly updates, shared outcome tags, sender-profile performance data, email infrastructure features, message and subject-line AI fields, scoring fields, and next-best-action labels.

The annual license is intended for internal model training, evaluation, GTM scoring, and workflow intelligence, with resale and redistribution excluded unless separately approved.

What's included
  • Larger privacy safe LinkedIn + email outcome dataset
  • Quarterly data updates with newly processed LinkedIn and email outcome records, refreshed sender-performance summaries, updated message and subject-line AI features, normalized outcome tags, scoring fields, and next-best-action labels
  • Shared outcome tags, scoring, and next-best-action labels
  • Sender-profile, email-infrastructure, message + subject AI fields
Discuss Annual License
Designed for
  • 01Internal AI-agent training
  • 02GTM scoring
  • 03Message testing
  • 04Sender-performance analysis
  • 05Reply classification
  • 06Outcome prediction
  • 07Next-best-action modeling
Privacy and Rights

Privacy-safe by design.

The buyer-facing dataset is structured around category-level fields, anonymized sender IDs, channel, message structure, outcome tags, scoring, and next-best-action labels.

Not included
  • Emails
  • Phone numbers
  • LinkedIn URLs
  • Company names
  • Company websites
  • Raw private reply text
  • Raw outbound message copy
  • Raw subject lines, unless separately approved
  • Identifiable sender names
Structured around
  • Category-level fields
  • Anonymized sender IDs
  • Channel
  • Message structure
  • Outcome tags
  • Scoring
  • Next-best-action labels

Evaluation and annual license terms restrict resale, redistribution, public release, and external sharing unless separately approved.

Sender profile fields are designed for aggregate sender-effect analysis, fairness testing, profile-performance research, and model evaluation. They are not intended for discriminatory targeting or automated decisions based on protected characteristics.

Get Access

Train and evaluate AI sales agents on real outreach outcomes.

Get access to the current dataset and quarterly updates of fresh outreach outcome intelligence.

shiras@smartreachai.com