Ground truth for AI agents that sell.
Not synthetic. Not a contact list. Real outreach outcomes mapped into a privacy safe schema.
15M+ outbound touches. 5 years. LinkedIn + email. Mapped to sender, message, CTA, reply, and outcome.
Explore Privacy-Safe GTM Training Data
Select a use case and see the type of structured sample output SmartReach can provide for AI agents, GTM engineering, revenue intelligence, and human-in-the-loop workflows.
13,504 booked calls analyzed.
Every booked call is tagged by industry, geography, company size, and decision-maker title, so your AI learns from outcomes that match your ICP.
The full outreach interaction · not the contact.
SmartReach AI captures who was targeted, what type of sender contacted them, what channel was used, how the message was structured, whether the CTA was hard or soft, whether a reply was received, what outcome occurred, and what action should happen next.
Each record is privacy safe and structured around lead category, sender profile category, channel, subject-line structure, message structure, CTA type, reply status, normalized outcome tag, scoring, and next-best-action labels.
Most GTM datasets tell you who a prospect is. This one shows what happened after outreach.
The sender profile layer is a key differentiator · buyers can study how sender credibility and profile signals move real campaign outcomes.
Real Campaign Outcomes
Records connect lead category, sender profile, message hook, CTA type, channel, reply status, and final outcome. AI agents can learn from what actually happened · not synthetic examples or raw contact records.
The Sender Profile Layer
Study how sender credibility, profile strength, region, follower band, connection band, age band, and gender or race/ethnicity category affect replies, scheduling, booked meetings, unsubscribes, and negative outcomes.
Multi-Channel Outcome Tags
LinkedIn workflow status and email response events unified under one normalized outcome vocabulary · booked, in talks, scheduling, not interested, unsubscribed, bounce, auto-reply, needs review, and more.
Not Contact Data
No emails, phone numbers, LinkedIn URLs, company names, or raw private reply text. Category-level fields and anonymized sender IDs · evaluate and license without exposing raw contact records.
Lead + sender + channel + hook + subject + message + CTA + outcome + next best action.
The full interaction · normalized into one schema across LinkedIn and email.
Six layers per record.
Every interaction is mapped into a shared six-layer schema where applicable.
Lead Intelligence
Seniority, role category, region, ICP segment, industry category, company size band, and buyer-role classification.
Sender Intelligence
Sender profile category, age band, gender category, race or ethnicity category, region, follower band, connection band, authority score, and profile strength.
Email Infrastructure
Sender type, email domain type, domain extension, ESP category, mailbox age band, domain reputation band, warmup status, and sending volume band.
Message Intelligence
Message hook, message length, tone, opening structure, question count, personalization level, CTA structure, value proposition type, and reading time.
Subject-Line Intelligence
Subject length, theme, structure, question flag, personalization flag, tone, specificity score, and reply-style signal.
Outcome Intelligence
Reply received, booked, scheduling, in talks, qualified, future interest, not interested, not relevant, unsubscribed, bounce, auto-reply, no-show, needs review, and next best action.
Explore the Email and LinkedIn data dictionaries.
Two privacy safe schemas. Field-by-field definitions, types, and examples. Toggle between datasets.
Built for teams shipping AI that sells.
Four buyer profiles. One dataset. Each one uses a different slice of the same 93+ field schema.
Train reply classifiers and next-message policies
Use 110K+ real human replies mapped to intent, sentiment, objection, and outcome. Fine-tune models that route, respond, and escalate the way top human SDRs do.
Fine-tune next best action models
Sender profile, message, CTA, reply, and booked call outcome in a single row. The exact schema needed for reinforcement learning on outbound decisions.
Benchmark outbound performance across industries
13.5K verified meeting outcomes across 5 years, segmented by industry, seniority, geography, and channel. A defensible baseline for your customers' dashboards.
Labeled B2B communication data at scale
Real outbound copy, real replies, real outcomes. Clean labels for supervised fine tuning, evaluation sets, and safety testing on business messaging.
Why nothing else gets you here.
The four paths a team usually considers before licensing. And what breaks on each one.
- ·Violates ToS and GDPR
- ·No outcome labels
- ·No sender or CTA context
- ·Contact records, not outcomes
- ·No reply text or intent labels
- ·Nothing to train a policy on
- ·No ground truth
- ·Model collapse on your own outputs
- ·Cannot evaluate against reality
- ·12 to 24 months to reach scale
- ·Single-industry bias
- ·No historical baseline
- First-party, consent-based collection
- 15M+ touches mapped to real outcomes
- 93+ labeled fields per row, both sides of the conversation
- 5 year historical baseline across industries
One real reply. Twelve labeled fields.
This is how a raw human response becomes a training row your model can actually learn from.
{
"input": {
"sender_role": "founder",
"prospect_seniority": "VP",
"message_variant": "intro_softCTA_v3",
"reply_text_tokens": [...]
},
"labels": {
"intent": "interested_needs_timing",
"objection": "budget_cycle",
"next_action": "offer_q1_pilot",
"outcome": "meeting_booked_delayed"
},
"reward": 0.72
}One shared taxonomy. Comparable answers.
Buyers can learn which lead types, sender profiles, channels, message hooks, subject-line structures, CTA types, and follow-up paths produce replies, scheduling, booked meetings, unsubscribes, bounces, not-interested responses, and not-relevant outcomes.
Compare LinkedIn and email under one outcome taxonomy · and test whether sender profile data, message structure, CTA type, and channel improve outcome prediction and next-best-action decisions.
LinkedIn Learnings
Which lead types respond, which sender profiles perform, which hooks work, and which workflow states should trigger follow-up, routing, or suppression.
Email Learnings
Which subject lines, sender types, domain types, ESPs, message structures, and CTAs produce replies, unsubscribes, bounces, or meeting intent.
Sender Learnings
Which sender profile categories produce replies, scheduling, booked meetings, not-relevant responses, or negative outcomes.
Cross-Channel Learnings
How LinkedIn and email compare by reply status, outcome tag, CTA type, message hook, sender type, and lead segment.
AI Decisioning Learnings
When an AI agent should follow up, wait, send a booking link, route to sales, suppress, mark not relevant, or request human review.
If you're building AI that sells across LinkedIn and email, this is your training data.
AI SDR Platforms
Train agents that pick the right sender, channel, hook, subject, and CTA · grounded in real multi-channel outcomes.
Sales Automation Tools
Ground your sequencing and next-best-action logic in normalized outcome tags across LinkedIn and email.
Outbound AI Agents
Learn the sender effect · how credibility, authority score, and profile band change reply and booking rates.
AI Model Evaluation
Test models against pre-derived message and subject-line features, including hook, tone, CTA structure, personalization level, and specificity.
Conversational AI
Classify replies: bounce, auto-reply, human, positive intent, clarification, negative · labeled at record level.
Revenue Intelligence
Enrich pipeline scoring with fit score, AI-agent fit, confidence, and next-best-action labels per record.
One normalized outcome tag across every channel.
LinkedIn replies and email replies get labeled with the same vocabulary · so a model can score, route, and evaluate them side by side.
Region, seniority, department, ICP segment, industry, company size band, redacted role category.
Anonymous sender key, authority score, follower + connection band, region, age band, replies + bookings produced.
LinkedIn or email · pending, accepted, awaiting reply, replied, failed, excluded · last action time.
Hook type, opening structure, tone, personalization level, CTA structure, subject theme, question flag, specificity.
Shared outcome tag across channels · outcome category · sales stage · outcome source.
Fit score, AI-agent fit, confidence, data-quality flag, next best action label.
One vocabulary means LinkedIn and email are directly comparable · for scoring, routing, evaluation, and next-best-action modeling.
Choose your refresh cadence.
New data is added every day. 10,000 email touches and 1,000 LinkedIn touches, labeled and mapped to the same schema. Choose how often that new data ships to you.
Each refresh gives buyers fresh examples of what is working now across channels, sender profiles, message structures, CTA types, reply patterns, unsubscribes, bounces, and booked-meeting outcomes.
New LinkedIn and email outcome records
Freshly processed outreach interactions mapped to the shared schema.
Updated outcome labels
Booked, scheduling, in talks, not interested, not relevant, unsubscribed, bounce, auto-reply, needs review, and other normalized tags.
Sender-performance refresh
Updated sender profile summaries, authority scores, profile strength bands, and outcome performance by sender type.
Message and subject-line updates
New message hooks, CTA structures, subject-line patterns, tone labels, personalization fields, and performance signals.
Fresh scoring fields
Updated fit scores, agent fit, confidence scores, data quality flags, and next-best-action labels.
Paid evaluation dataset.
A structured package to test ingestion, outcome classification, sender-performance signals, message and subject-line features, next-best-action logic, and model lift · before committing to the full license.
The evaluation package is designed for internal testing, schema review, model benchmarking, outcome classification, and next-best-action analysis. Production use, resale, redistribution, and public release require a separate annual license.
Request Evaluation Dataset- 01CSV and JSONL formats
- 02Shared outcome taxonomy
- 03Data dictionaries
- 04Sender profile summary
- 05Message and subject-line AI feature guide
- 06Field coverage report
- 07Evaluation terms
The full dataset, for a year.
The standard annual license provides access to the larger privacy safe LinkedIn and email outcome dataset · with quarterly updates, shared outcome tags, sender-profile performance data, email infrastructure features, message and subject-line AI fields, scoring fields, and next-best-action labels.
The annual license is intended for internal model training, evaluation, GTM scoring, and workflow intelligence, with resale and redistribution excluded unless separately approved.
- Larger privacy safe LinkedIn + email outcome dataset
- Quarterly data updates with newly processed LinkedIn and email outcome records, refreshed sender-performance summaries, updated message and subject-line AI features, normalized outcome tags, scoring fields, and next-best-action labels
- Shared outcome tags, scoring, and next-best-action labels
- Sender-profile, email-infrastructure, message + subject AI fields
- 01Internal AI-agent training
- 02GTM scoring
- 03Message testing
- 04Sender-performance analysis
- 05Reply classification
- 06Outcome prediction
- 07Next-best-action modeling
Privacy-safe by design.
The buyer-facing dataset is structured around category-level fields, anonymized sender IDs, channel, message structure, outcome tags, scoring, and next-best-action labels.
- Emails
- Phone numbers
- LinkedIn URLs
- Company names
- Company websites
- Raw private reply text
- Raw outbound message copy
- Raw subject lines, unless separately approved
- Identifiable sender names
- Category-level fields
- Anonymized sender IDs
- Channel
- Message structure
- Outcome tags
- Scoring
- Next-best-action labels
Evaluation and annual license terms restrict resale, redistribution, public release, and external sharing unless separately approved.
Sender profile fields are designed for aggregate sender-effect analysis, fairness testing, profile-performance research, and model evaluation. They are not intended for discriminatory targeting or automated decisions based on protected characteristics.
Train and evaluate AI sales agents on real outreach outcomes.
Get access to the current dataset and quarterly updates of fresh outreach outcome intelligence.
shiras@smartreachai.com