Analysis 9 min read

Your Meeting Data Is the Next AI Training Goldmine

AI notetakers hold thousands of hours of your company's conversations. Who owns that data, who trains on it, and what the privacy policies say.

meetingstack research ยท 9 min read

Updated October 3, 2026: we re-checked every vendor policy claim against the live privacy, pricing and trust pages. The original version wrongly said Granola's training opt-out was Enterprise-only and that Otter keeps data after deletion for "legitimate business purposes." We also corrected HIPAA availability, removed claims we could not source (including an estimated-volume chart), and added links. Details are on our corrections page.

Take a 50-person company that records 15 one-hour meetings a day. Over 21 working days, that is about 300 hours of recorded audio per month (our illustration, not a measured figure). That is 300 hours of unfiltered conversation about product strategy, deal pricing, employee performance, legal exposure, and competitive intelligence. Now multiply that across every company feeding its meetings into AI notetakers. The volume of sensitive data flowing through these tools is large, and most teams have never read the privacy policy.

We read six vendors' privacy policies and trust pages so you do not have to (re-checked October 3, 2026). What we found ranges from strict to permissive.

Meeting recordings are not like email or Slack messages. They capture tone, hesitation, who talks over whom, who stays silent. They contain the things people say when they think only their colleagues are listening. For AI companies, this data is extraordinarily valuable. For the companies producing it, the risk is just as large.

What your notetaker knows about you

The transcript is the obvious output. But the raw data captured by AI notetakers goes far deeper than words on a page.

Voice biometrics. Every recording contains a unique voiceprint for each speaker. Voice data can identify individuals across meetings, even across different tools. Once captured, it cannot be changed like a password.

Organizational maps. Meeting metadata reveals who talks to whom, how often, and for how long. Over weeks, a notetaker builds a precise map of your org chart, decision-making hierarchy, and internal alliances. No employee directory is this accurate.

Deal intelligence. Sales calls contain pricing, discount thresholds, competitor comparisons, and objection patterns. A single quarter of recorded sales meetings is a complete playbook for how your company sells. As we explored in our Gong Tax breakdown, revenue intelligence tools ingest this data at scale.

Product roadmaps. Internal planning meetings expose upcoming features, timelines, technical debt, and strategic priorities. This is the information competitors pay consultants to uncover.

HR and legal exposure. Performance reviews, disciplinary discussions, termination meetings, legal strategy sessions. These get recorded too, sometimes intentionally, sometimes because someone forgot to pause the bot. The liability is enormous.

This is not just text. It is a living, searchable map of how your company operates, makes decisions, and manages its people. Every AI notetaker that processes your meetings holds a copy.

What the privacy policies actually say

We read the privacy policies and security documentation of six major AI meeting tools. The differences are significant. Here is what we found, with each policy's own date in the text below:

Tool Trains on data? Opt-out? Data residency Retention SOC 2 HIPAA
Otter.ai Yes (de-identified) None found in policy US (AWS) "As long as necessary"; no fixed timeline Type II Enterprise add-on
Fireflies.ai No N/A (no training) US and other countries; private storage on Enterprise Zero retention by vendors; account data deleted 30 days after closure Type II Enterprise
Fathom De-identified* Yes (account settings) US User-controlled; account deletion within 30 days Type II Enterprise (BAA)
tl;dv No N/A (no training) Primarily EU (EEA) User-controlled Yes (type not stated) Not listed
Gong No N/A (no training) Not verified Up to 3 years or contract term; deleted 30 days after termination Type II HIPAA controls mapped in SOC 2; no BAA stated
Granola Yes (de-identified) Yes, per user on every plan; Enterprise off by default US (AWS) Audio deleted after transcription; free plan shows 30 days of history Type II Enterprise (BAA)
Green = good for users. Red = concern. Amber = conditional or limited. *Fathom uses de-identified data to train its in-house models; users can opt out in account settings.

Based on our review of publicly available privacy policies, pricing pages and trust pages, checked October 3, 2026. "Not verified" means we could not confirm the claim on a vendor page. Policies change frequently. Verify directly with each vendor before making decisions.

A few things stand out. Otter.ai's privacy policy (effective June 16, 2026) lists "training our proprietary AI technology on de-identified audio recordings and on transcriptions (which may contain Personal Information)" among its uses of your data, and we found no training opt-out in it. On retention, Otter keeps personal information "for as long as necessary" for the policy's purposes or as the law requires, and says deleted data is made "irrecoverable." Otter also faces a consolidated class action over recording consent; on August 13, 2026, a federal court in California let the core wiretap and privacy claims proceed. No court has found that Otter broke the law.

Fireflies takes the opposite approach. Its privacy policy (last updated September 28, 2026) says meeting content falls under a "Zero Data Retention policy": third-party vendors do not store it after processing, and it is not "used for training internal or external AI models." Enterprise plans add private storage and custom data retention, per its pricing page.

Gong, one of the largest players in revenue intelligence, states plainly on its trust page that customer data "is never used to train generative models."

Granola is the newest entry, and its defaults are the thing to check. Its privacy policy (effective September 8, 2026) says it uses only de-identified data to train AI models, "which you can opt-out of within your Granola account settings." That opt-out is per user on every plan, including free, per Granola's pricing page; Enterprise workspaces are opted out by default and admins can enforce it. Notes can also be shared by link: on Free and Business plans, each user picks a default of "Anyone with the link", "Only my company" or "Private", according to Granola's sharing docs. For teams sharing sensitive meeting content, both settings deserve a look on day one.

The volume matters. The largest notetakers operate at scale: Fireflies' pricing page says it is used across more than 1 million companies, and Fathom reported over 400,000 monthly active users when Superhuman acquired it in September 2026. Even if only a fraction of that audio touches model training pipelines, the corpus is large.

The training data question

The central question is simple: is your meeting data training someone else's AI?

Several vendors say no, with different scope. Fireflies says meeting content trains no AI models; Gong and tl;dv say it does not train generative models. Ask for the data processing agreement to see whether the contract says the same.

Free tiers are a different story. When a tool costs nothing, the business model has to come from somewhere. Otter's policy covers training on "de-identified" audio and transcripts, with no plan-level exception and no opt-out that we found. Granola trains on de-identified data by default unless each user switches it off (Enterprise workspaces start opted out). Fathom does the same with its in-house models, also with an opt-out. The word "de-identified" does heavy lifting in these policies, and it is fair to ask how a vendor de-identifies audio that still carries a voiceprint.

There is also the subprocessor question. Even tools that do not train their own models send your audio to third-party transcription and LLM providers. Fireflies says its vendors do not store meeting content after processing or train on it. tl;dv's privacy policy says it may use Anthropic models through Google Cloud Vertex AI, and that its AI providers do not use customer content to train their models. Fathom says it does not authorize third parties such as OpenAI, Anthropic or Google to train on meeting content.

But not every vendor is this careful. Smaller tools may use default API configurations where the LLM provider retains inputs for 30 days or uses them for abuse monitoring. Unless the notetaker vendor has negotiated specific terms, your meeting data may sit on OpenAI or Anthropic servers longer than you expect.

The risk compounds over time. A single meeting transcript is moderately sensitive. A year of transcripts from every meeting in your company is a complete intelligence file. The longer data is retained, and the more broadly it is shared, the larger the attack surface becomes.

What to ask before you sign up

Before adopting any AI notetaker, your security team should get clear answers to these questions:

  • Does any of our meeting data (audio, transcripts, metadata) enter training pipelines? "De-identified" is not the same as "no." Press for specifics on what de-identification means and whether it applies to audio or only text.
  • Which subprocessors handle our data, and what are their retention terms? A tool may not train on your data, but its transcription provider might retain it. Ask for the subprocessor list and the DPA for each.
  • Where is our data stored, and can we choose the region? If you operate under GDPR, data residency is not optional. Some tools offer EU hosting; others store everything in US-East regardless.
  • What happens to our data when we cancel? Fathom and Fireflies both cite a 30-day deletion window after an account deletion request or closure. Gong deletes all data 30 days after contract termination. Otter's policy says only that it keeps data "for as long as necessary." The difference matters.
  • Can we get a BAA for HIPAA compliance? If your organization handles any health-related discussions (benefits, insurance, patient data), you need this. Not every tool offers it.
  • Is there an audit log for who accessed our recordings? If an employee shares a recording externally, or if a vendor engineer accesses it during a support ticket, you should know.
  • What is the incident response plan if our data is breached? SOC 2 certification means a company has controls in place. It does not guarantee those controls will hold. Ask about breach notification timelines and past incidents.
  • Can your vendor reconstruct who attended which meetings with whom, even after you delete transcripts? Metadata (participant lists, timestamps, calendar links) often persists long after content is removed. This is the data that maps your organization.

If your vendor cannot answer these questions clearly, in writing, that tells you something.

Who gets this right

Several tools stand out for handling data responsibly.

Fathom has built its brand on privacy. Its privacy policy (last updated August 16, 2026) bars third parties from training on meeting content, and users can opt out of Fathom's own de-identified training in account settings. It reports SOC 2 Type II and offers a signed HIPAA BAA on Enterprise, per its pricing page. For a free notetaker, the privacy posture is unusually strong. Two caveats. Fathom's free tier lacks the CRM field sync that mid-market teams need; that sits on the Business plan ($34 monthly, $25 annual). And Superhuman acquired Fathom on September 14, 2026; Fathom says nothing changes for now, but future policy will come from a new owner.

Fireflies goes further on the infrastructure side. Enterprise customers get private storage and custom data retention. Third-party vendors operate under its zero data retention policy. Its pricing page lists SOC 2 Type II and GDPR compliance, with HIPAA compliance on Enterprise.

Gong takes a conservative public position on AI training: its trust page says customer data is never used to train generative models, with no plan-tier exception stated. For companies processing high-stakes sales conversations, this matters.

Bluedot avoids the visible bot. It records through a Chrome extension or desktop app instead of joining as a participant, per its site. Recordings still go to Bluedot's cloud for processing, so this removes the bot, not the vendor. Its security page says customer data is never used to train AI models, lists SOC 2 Type II and GDPR controls, and says it supports EU data residency.

Krisp is often described as an on-device tool, but that does not hold for its notetaker. Krisp's security page says its AI Meeting Assistant stores meeting transcripts and recordings in Krisp Cloud, and its privacy policy says it shares meeting content with third-party providers to generate AI summaries. Organizations that cannot tolerate any data leaving their network should not treat it as a local-only option.

On the other end of the spectrum, Otter trains on de-identified user data with no opt-out in its policy. Granola and Fathom also train on de-identified data by default, but each offers every user an opt-out in settings. The difference between "no opt-out" and "opt-out you have to find" matters more than any marketing claim.

The pattern is mixed. Price does not predict data practices: Fathom's and Granola's policies apply the same training terms, with the same opt-out, to free and paid individual plans, and Otter's applies its training terms to both with no opt-out. Meeting recordings are too sensitive to assume either way. If your notetaker is free, read the training and retention sections of the privacy policy before you assume your conversations are not the product.