TUESDAY, SEPTEMBER 29, 2026|No. 16913
Technology · Privacy

AI Chatbots Share Sensitive User Data with Advertisers, Study Finds

A new study reveals that popular AI chatbot services are disclosing sensitive user conversation data, including prompts and screenshots, to third-party advertisers.

A digital representation of artificial intelligence interacting with data streams.
A digital representation of artificial intelligence interacting with data streams. · Photo by Steve A Johnson on Unsplash
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents

Guilherme Oliveira IMDEA Networks

Roi S. Serna IMDEA Networks/UC3M

Aniketh Girish IMDEA Networks

Miguel Sanchez IMDEA Networks

Tautvydas Jackevicius IMDEA Networks

Guillermo Suarez-Tangil IMDEA Networks

Juan Manuel De Santa Olalla Gómez IMDEA Networks

Jorge Garcia-Herrero Independent

Narseo Vallina-Rodriguez IMDEA Networks

Abstract

As prominent conversational AI providers like OpenAI adopt advertising-based business models, traditional web and mobile tracking practices are expanding into conversational AI services [15]. However, despite their growing adoption, the tracking, data-sharing, and monetization practices of conversational AI services remain largely opaque and have received comparatively limited scrutiny from researchers, regulators, and the public.

In this paper, we present a systematic privacy analysis of the web and mobile deployments of nine prominent conversational AI services. Using a combination of static and dynamic analysis, we study the presence of third-party Advertising and Tracking Services (ATSes), characterize their data flows, and evaluate how consent choices, subscription tiers, and access-control mechanisms influence conversation exposure to third parties. We uncover privacy risks unique to conversational AI platforms: multiple providers disclose sensitive conversation-derived artifacts—including titles, prompts, and screenshots—to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation.

Our findings reveal how traditional tracking technologies are increasingly intertwined with AI-mediated interactions, creating new pathways through which sensitive user and conversational information can be collected, inferred, and disseminated. To assess the broader implications of these practices, we analyze them in the context of the GDPR and ePrivacy Directive. We conducted a responsible disclosure process involving affected providers and competent European Data Protection Authorities. Our results demonstrate that conversational AI services introduce a novel privacy attack surface in which provider-generated conversational artifacts become subject to tracking and public exposure, highlighting the need for stronger safeguards governing AI-mediated interactions.

Keywords

Conversational AI, LLMs, Privacy, Mobile, Web, Trackers

1 Introduction

Recent advances in Large Language Models (LLMs) have enabled the emergence of conversational AI services such as ChatGPT, Gemini, and Claude, capable of supporting persistent interactions, multimodal processing, and autonomous task execution. As adoption of these services grows for personal and professional activities [43], service providers are exploring new business models to monetize their growing user bases.

Advertising is emerging as one such model, potentially extending into conversational AI the tracking and attribution infrastructures traditionally associated with web and mobile platforms. For example, Reuters reported that OpenAI partnered with Criteo to conduct an advertising pilot for ChatGPT free-tier users in the United States in early 2026 [52].

However, the integration of these tracking technologies raises distinct privacy concerns. Unlike traditional web and mobile applications, conversational AI services routinely process highly sensitive prompts, contextual information, behavioral patterns, uploaded documents, and persistent interaction histories that may reveal intimate aspects of users’ lives and professional activities. The disclosure of such information to third-party tracking services, particularly without meaningful transparency or consent, may therefore expose users and organizations to significant privacy risks.

Prior work by Jazlan et al. has examined the integration of thirdparty tracking in web-based conversational AI services [32], primarily focusing on identifying trackers and characterizing their data collection practices. However, the unique interaction models of conversational AI services introduce new privacy risks across their web and mobile clients: these services generate conversationderived artifacts—including conversation identifiers, URLs, titles, previews, prompts, responses, and interaction metadata—that may be disclosed to third parties or exposed through publicly accessible resources. Moreover, how these exposures are shaped by by consent choices, privacy settings, subscription tiers, and access-control mechanisms remains largely unexplored. To address this gap, we investigate three research questions:

• RQ1: To what extent do conversational AI services integrate third-party tracking, analytics, advertising, and attribution infrastructures across their web and mobile clients?

This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this license visit https://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA. Proceedings on Privacy Enhancing Technologies YYYY(X), 1–18 © YYYY Copyright held by the owner/author(s). https://doi.org/XXXXXXX.XXXXXXX

• RQ2: What conversation-derived artifacts and user information are exposed by conversational AI services, either to third-party entities or through publicly accessible resources, and what privacy risks emerge from their disclosure?


• RQ3: How do cookie consent choices, subscription tiers, privacy settings, and access-control mechanisms shape the disclosure and accessibility of information in conversational AI services?

To answer these questions, we conduct a systematic privacy analysis of nine prominent conversational AI services, covering the web clients of all nine providers and the Android clients of the eight that offer an Android mobile app. We combine static and dynamic analysis to evaluate their privacy practices across consent choices, subscription tiers, and access-control configurations. Specifically, we make the following contributions:

(1) Across the evaluated services, we identify 44 third-party organizations and observe that every evaluated AI service integrates at least one third-party advertising or tracking service. We further uncover substantial differences between web and Android clients and identify third-party services that are activated only after users explicitly accept non-essential cookies, demonstrating that consent decisions directly influence the tracking surface of conversational AI platforms (§5).

(2) We uncover novel privacy risks specific to conversational AI services. Unlike traditional tracking systems that primarily observe browsing activity, conversational AI platforms generate artifacts that directly encode user interactions. We show that 6/9 web and 3/8 Android clients disclose conversation URLs, titles, prompts, and screenshots to third-party services, often alongside persistent user identifiers. We further demonstrate that the privacy implications of these disclosures are strongly shaped by consent choices and sharing functionality, with several providers exposing entire conversations through publicly accessible permalinks lacking access controls. These findings reveal new channels through which trackers and external actors can gain access to users’ entire conversations (§6).

(3) We conduct a legal analysis of observed practices under EU data-protection law, assessing the compatibility of tracker activation, consent mechanisms, conversation-artifact disclosures, identity-linkage practices, and publicly accessible conversational resources with the GDPR and ePrivacy Directive (§7).

Our findings show that integrating traditional tracking infrastructures into conversational AI services creates novel pathways for exposing sensitive user information, including conversationderived artifacts and publicly accessible conversations. More broadly, our results challenge the perception of conversational AI services as confidential exchanges between users and AI providers. Instead, they are increasingly integrated into the broader online tracking ecosystem, raising important technical and regulatory challenges for AI-mediated services.

Responsible Disclosure. We followed a responsible disclosure process for all identified issues, notifying affected providers and the competent Data Protection Authorities (DPAs) as described in the Ethical Considerations section.

2 Background

This section provides background on conversational AI services and their growing integration with tracking technologies (§2.1), and on tracking mechanisms commonly deployed across web and mobile platforms (§2.2).

2.1 Conversational Agents

Conversational AI services are LLM-based systems that interact with users through natural language interfaces, typically accessible through web and native mobile clients. The public release of ChatGPT by OpenAI in November 2022 marked a major inflection point in the AI industry, rapidly reaching hundreds of millions of users and triggering an industry race to deploy conversational AI platforms across consumer and enterprise ecosystems [42]. Since then, other providers have released competing services, including Perplexity AI, Anthropic’s Claude, Google’s Gemini, Microsoft’s Copilot, xAI’s Grok and DeepSeek.

Despite differences in architecture and deployment models, all these services share several common characteristics. Most conversational AI services maintain persistent user accounts and interaction histories, and integrate external services such as search engines, analytics platforms, telemetry frameworks, advertising infrastructures, and cloud-hosted APIs. Modern AI agents also increasingly support multimodal capabilities, including image analysis, voice interaction, document analysis, browsing assistance, and autonomous task execution.

The rapid development and adoption of these services amplify the economic incentives to introduce data-driven monetization models. Recent industry developments and press releases suggest that conversational AI services are beginning to integrate into the existing advertising and tracking ecosystem rather than replacing it. For example, Criteo reported that 40% of surveyed U.S. consumers already use AI agents for product discovery and shopping assistance [15], while industry actors increasingly describe agentic AI as the next opportunity for targeted advertising and “agentic commerce”_ [2]. Similarly, major tech companies are developing infrastructure that enables AI agents to interact directly with commercial platforms and external digital services, such as Google’s Universal Commerce Protocol (UCP), which facilitates AI-driven commerce and interoperable agent ecosystems [4].

2.2 Mobile and Web Tracking

Modern web and mobile services are deeply intertwined with thirdparty advertising and tracking services that enable user profiling, personalization, attribution, and targeted advertising at scale [48]. On the web, trackers rely on techniques such as cookies, tracking pixels, browser fingerprinting and cookie syncing to collect browser metadata, interaction events, device characteristics, network information, and account-linked identifiers, enabling persistent crosssite identification and behavioral profiling [1, 17, 29, 34]. Mobile applications similarly integrate third-party SDKs [24, 48] that collect device and behavioral data, frequently relying on platformsupported identifiers such as the Android Advertising ID (AAID) and Apple’s Identifier for Advertisers (IDFA) [28, 53], and hashed email addresses (HEMs)1 to support cross-device tracking [62, 64].

Both browsers and mobile operating systems provide mechanisms to limit such tracking. Web browsers increasingly restrict third-party cookies, fingerprinting, and other cross-site tracking

1Hashed email addresses (HEMs) are pseudonymous identifiers derived from users’ email addresses that can enable identity matching across services without transmitting the email address in plaintext [25].


Conversational AI service

Client-Side Android Web Server-Side Backend Agent

Prompt Response Embedded 3rd-Party SDK

User

ATS Server

CA Conversational Artifacts

User Identifiers UserID, DeviceID, AnonID, OrganizationID, HEM

Conversation identifiers Convld, ChatURL, SharedID, ShareURL

Chat Content Title, prompt, screenshot

Figure 1: Privacy risks.

techniques, while content blockers and privacy-enhancing extensions can prevent requests to known tracking domains. Mobile platforms instead rely on application sandboxing, runtime permissions, and restrictions on access to advertising identifiers and other sensitive resources. However, these protections do not eliminate third-party data collection: embedded SDKs execute within the host application’s security context and may access data and permissions available to the application [28]. These differences in tracking mechanisms and platform protections motivate our separate analysis of the web and mobile clients of conversational AI services, as explained in §4.

3 Privacy Risks in Conversational AI Services

Integrating traditional advertising and tracking technologies into conversational AI services introduces unique privacy risks. Unlike conventional web and mobile services, these services routinely process highly sensitive and contextual information, including freeform conversations, behavioral interactions, uploaded documents, interaction histories, and persistent user profiles. Consequently, embedded third-party tracking technologies can access not only behavioral and device information, but also artifacts that encode the content and context of users’ interactions with the AI providers.

We consider three primary entities, shown in Figure 1: the user, the conversational AI service (the first party), and third-party ATSes such as analytics providers, advertising networks, crash reporting services, and anti-fraud products. The conversational AI service comprises both client- and server-side components. Third-party code and libraries embedded in web or mobile clients may collect and transmit information directly from the clients to their cloud infrastructure, while server-side components may disclose information to third parties outside the boundaries of user devices.

Within this setting, new privacy risks emerge when conversationderived artifacts are disclosed to third parties. These artifacts may include prompts containing information directly provided by users, conversation titles that summarize the content of an interaction, screenshots capturing conversational context, or persistent conversation URLs (permalinks). Unlike tracking methods on conventional web and mobile services, conversation artifacts can directly reveal the content, context, and potentially sensitive nature of users’ interactions with the AI services, in addition to the exposure of model responses to potential competitors.

1 Conversational AI Service Selection 9 Services Selected ChatGPT Claude Grok DeepSeek Gemini Copilot Perplexity Le Chat MS Copilot Meta AI $4.1

2 Instrumentation & Black Box Analysis Methods Website Analysis Chrome 148 / DevTools (CDP) HTTP/S Requests Cookies Javascript $4.2

Android App analysis Pixel 3a / Android 12 STATIC: Androguard DYNAMIC: Instrumented AOSP $4.3

3rd Party Classification eTLD+1 1st party 3rd party ATS 3P API

4 Analysis Dataflows and Leaked Conversation Artifacts Third-party analysis $5 Privacy analysis $6

Consent Forms and subscription tiers tier (guest-free-paid) consent (ignore-reject-accept) mode(default-incognito) $4.4

Conversation and Prompt inputs Sensitive information I have "X" condition AI Information .... $4.5

3 Experiment Conditions and Input

Figure 2: Methodology overview.

The risks are amplified when such information is disclosed alongside persistent user or device identifiers like email hashes, allowing third parties to associate sensitive conversational data with individual users and potentially link it across sessions or services. Moreover, conversation URLs may provide access to additional content when the underlying resources lack adequate access controls, potentially exposing the entire conversation to third parties.

These risks challenge the perception of conversational AI as a confidential interaction between a user and an AI provider. When combined with conventional tracking infrastructures, the rich and persistent artifacts generated by conversational AI create new channels through which sensitive information can be disclosed, linked to individual users, or made accessible to third parties.

4 Methodology

Figure 2 provides an overview of our research methodology to answer our three research questions. Following this workflow, we first describe the selection of representative conversational AI services (§4.1), followed by our instrumentation and black-box analysis methods for their web (§4.2) and Android clients (§4.3), how we assess the effects of consent choices and subscription tiers (§4.4), and the conversation inputs used to trigger systematic and reproducible behaviors on the conversational AI services (§4.5). All the experiments were conducted in Spain during May 2026.

4.1 Conversational AI Service Selection

We analyze a set of prominent conversational AI services supporting both web-based and Android-based clients. Rather than maximizing breadth, we conduct an in-depth and systematic analysis on a small set of representative AI services that account for a substantial share of the market. Due to the lack of reliable market share figures per provider, we select those with large user bases using objective popularity proxies, including Tranco rankings [35] for web services and cumulative Google Play installation counts for mobile.

Table 1 summarizes the nine services we include in our study together with their providers, web and mobile implementations, and the popularity indicators that guided our selection. All these


Table 1: Web and mobile conversational AI services analyzed in this study ordered by Tranco rank domain (May 2026).

ServiceProviderWeb DomainAndroid PackageTranco RankPlay Store Installs
ChatGPTOpenAIchatgpt.comcom.openai.chatgpt48>1B
ClaudeAnthropicclaude.aicom.anthropic.claude617>10M
GrokxAIgrok.comai.x.grok956>100M
DeepSeekDeepSeekchat.deepseek.comcom.deepseek.chat1,196>50M
PerplexityPerplexity AIperplexity.aiai.perplexity.app.android1,249>100M
GeminiGooglegemini.google.comcom.google.android.apps.bard6,045>1B
MS CopilotMicrosoftcopilot.comcom.microsoft.copilot11,011>50M
Mistral (Le Chat)Mistral AIchat.mistral.aiai.mistral.chat12,039>1M
Meta AIMetameta.aicom.facebook.stella13,701>50M

services remain in the top-14K services for Tranco2 in May 2026. Three of them are in the Top-1K: ChatGPT (top-48), Claude (top- 617), and Grok (top-956). On the mobile side, all their mobile apps have at least 1M cumulative installs, and two of them (Gemini and ChatGPT) have more than 1B cumulative installs.

4.2 Website Analysis

We analyze the presence of trackers and their data collection practices on web-based conversational AI services using Google Chrome (v148.0.7778.167) Developer Tools (CDP), and store the resulting HAR traces for subsequent analysis. This setup enables observation of HTTP(S) requests, JavaScript execution, browser storage access, cookies, tracking pixels, and other client-side tracking mechanisms that are generated during user interactions.

We then inspect the communication to endpoints including request parameters, request bodies, cookies, browser storage entries, and protocol metadata to identify the transmission of (i) user identifiers (e.g., account IDs and email addresses); and (ii) conversationspecific metadata, conversation content, and identifiers (e.g., chat identifiers and sharing links). We also search for transformed representations of such data, including Base64 encodings and common hashing algorithms used to generate hashed email addresses (SHA- 256, SHA-1, and MD5). We inspect browser storage mechanisms and monitor stable IDs across sessions to identify persistent identifiers and tracking artifacts. Finally, all observed disclosures are mapped to their corresponding recipient domains and correlated with the experimental configuration that triggers them.

All sessions are conducted manually by a researcher and cover authentication, onboarding, and pre-defined conversational exchanges, as further developed in §4.5. Experiments are repeated across subscription tiers, platform configurations, and privacy conditions, as detailed in §4.4, to evaluate the impact of consent scenarios (acceptance or rejection of non-essential cookies) and subscription tiers (guest, free, and premium accounts). Pilot experiments show highly deterministic tracking behavior for every configuration, and therefore each configuration is analyzed once.

Third-party Domain Classification. Labeling of tracking domains was performed manually by a co-author with over a decade of research experience, applying conservative criteria to avoid overreporting. We distinguish between first- and third-party domains using publicly available blocklists and tracker intelligence sources,

including uBlock Origin [59] and whotracks.me [9] with the support of DNS lookups and Certificate Transparency logs [7]. Because some conversational AI providers also operate advertising, analytics, and cloud infrastructures (e.g., Google, Microsoft and Meta), considering corporate ownership alone may obscure tracking relationships. We therefore classify provider-owned advertising, analytics, and telemetry endpoints as third-party Advertising and Tracking Services (ATSes). These services may facilitate data sharing with other ad-tech actors through Real-Time Bidding requests, or even operate under separate legal entities, as in the case of xAI and X Corp. This approach provides a consistent comparison of tracking practices across all services.

4.3 Android App Analysis

We analyze conversational AI Android apps through a combination of static and dynamic techniques to maximize behavior coverage:

Static Analysis. We decompile every app’s APK using Androguard [16]. We parse AndroidManifest.xml to enumerate declared sensitive permissions relevant to tracking ( e.g., AD_ID, READ_PHO- NE_STATE, ACCESS_FINE_LOCATION). We identify embedded thirdparty SDKs by extracting package namespaces and mapping the app’s package name (e.g., com.company.app) to its corresponding eTLD+1 (company.com) following prior work practices [24, 28, 64]. Packages and contacted domains whose ownership does not match the app’s eTLD+1 are treated as third-party components [48, 53]. We then manually match these packages and domains using public SDK documentation and prior work mappings [28, 45], and the third-party classification method described in §4.2.

Dynamic Analysis. We execute each app on an instrumented Google Pixel 3a using an Android 12 build that transparently monitors runtime access to permission-protected APIs, file I/O operations, and all outbound network traffic; equivalent coverage is achievable using mitmproxy [10] and Frida [46]. We observe reads and writes to TLS sockets at the system level, enabling traffic inspection without certificate injection and without disrupting TLS handshakes, including in certificate-pinned apps [44, 47]. The instrumentation traces access to sensitive resources including device IDs (AAID, Android ID, IMEI, GSF ID, Boot ID), hardware IDs (WiFi MAC address), network scan data (WiFi SSIDs, BSSIDs), and account-linked IDs ( e.g., email address).3 To complement static

2 https://tranco-list.eu/list/3Q25L/1000000

3Each device is provisioned with pseudonymous IDs (email address, phone number) to register test accounts on each platform. Because ID values are known per device,


SDK detection, we instrument the Android Runtime to log classes loaded at runtime by tracking the FindClass method of the class linker, enabling identification of obfuscated SDK components that static analysis may miss. Captured traffic is automatically decoded for common encodings (gzip, Base64) and SDK-specific obfuscation methods, and parsed to extract field names, values, and destination endpoints using the third-party endpoint classification method described in §4.2. Sessions are screen-recorded to support post-hoc verification of data flows.

4.4 Consent Forms and Subscription Tiers

We conduct controlled experiments across different consent choices and subscription tiers, allowing us to assess how these factors shape the observed privacy risks. Our preliminary experiments revealed differences between web and mobile consent mechanisms. While web services generally present cookie consent banners, most Android apps do not expose an equivalent consent mechanism at launch. We therefore evaluate the consent conditions and privacy settings supported by each client.

For web clients, we perform separate runs in which we explicitly accept or reject non-essential cookies through the consent mechanisms offered by each service. These interactions are performed manually rather than through automated means to accurately capture service-specific consent flows and privacy controls. Each experiment starts from a fresh browser profile, with all local state removed to prevent previously stored identifiers, consent choices, authentication tokens, or telemetry artifacts from influencing subsequent measurements.

We additionally conduct controlled interactions across guest, free, and paid subscription tiers when supported by each service. For authenticated tiers, we maintain separate accounts and execute equivalent interaction scenarios across configurations. Meta AI and DeepSeek do not offer premium tiers, and neither provides a guest tier. For guest and paid configurations, we accept all cookies to capture flows to third-party organizations and isolate differences associated with authentication and subscription modes.

4.5 Conversation and Prompt Inputs

To capture dynamic evidence of the dissemination of conversationderived artifacts to third-party services alongside user identifiers, we manually interact with every service on each platform following a predefined and reproducible interaction protocol. For each experimental mode described in §4.4, we conduct a chat session comprising several prompts designed to trigger a broader range of behaviors. When supported by the service, we additionally generate and open a chat-sharing link in a separate browser session to assess whether overly permissive access controls expose shared conversations to third parties.

While exhaustive coverage of all interaction paths is unfeasible, the prompts were iteratively refined during a preliminary measurement campaign to capture representative behaviors. Rather than aiming for absolute completeness, we use rich, health-related content to emulate realistic users following a persona-based auditing approach (§9). Specifically, the prompts emulate a user consulting

Table 2: Most frequent third-party organizations across the nine conversational AI services tested. We report their presence separately for web and mobile clients, as well as across both client types (∩) and either client type (∪). Entries are sorted by the number of services in which the organization is present on both web and mobile clients. Organizations are broken down by product when possible. Legend: G# web only, H# mobile only, both web and mobile, and empty if not contacted.

Organization / product# ServicesPer service
Organization / productWebMobileChatGPTClaudeCopilotDeepSeekGeminiGrokMeta AIMistralPerplexity
Organization / productWebMobileChatGPTClaudeCopilotDeepSeekGeminiGrokMeta AIMistralPerplexity
Google8879●●●●○●○●●
Firebase0808○○○○○○○○
Search8228○○○○○○●●
Ads7007○○○○○○○
Tag Manager7007○○○○○○○
Accounts5005○○○○○
Sentry2424○○●●
Datadog3223●●○
Intercom2103○○○
Meta3113○○●

the service about a medical condition, introducing sensitive context to evaluate platform behavior under realistic and privacy-sensitive scenarios. We discuss the limitations of this targeted approach in §8.1; nevertheless, it provides a consistent baseline for comparing behaviors across conversational AI services.

5 Third-party Service Analysis

We study the structural integration of third-party services across the web and mobile clients of conversational AI services (§5.1) and how consent choices and subscription tiers influence their presence (§5.2). We then characterize the data disseminated to third-party organizations in §6.

5.1 Web vs. Mobile Tracking

Every conversational AI service contacts at least one third-party organization categorized as an ATS. Across our measurements, we observe 124 distinct third-party domains, which we attribute to 44 organizations, of which 34 are ATSes. Figure 3 captures the ATS ecosystem observed across services and client types.

However, third-party integration differs substantially between web and mobile clients. While 11 ATSes appear on both client types, 15 are observed exclusively on the web, including Google Tag Manager, TikTok, and consent-management platforms such as OneTrust. In contrast, 8 ATSes are found exclusively on mobile clients, including Braze.

Google products are the most pervasive across client types, as reported in Table 2, followed by Sentry, Meta, Datadog, and Intercom. The presence of these organizations reflects the broad reliance of conversational AI services on third-party products and services for functions including error monitoring and observability, customer support and engagement, but also advertising and analytics. Google

we can automatically search captured traffic for direct occurrences and their common hash transformations (MD5, SHA-1, SHA-256).


Google Search Grok Google Ads Google Tag ManagerGoogle Firebase Google AccountsPerplexity Sentry Claude Datadog Google APIs Meta Google User ContentChatGPT RevenueCat Perplexity Intercom Datadog SentryGoogle Search AppleClaude CopilotAdjust Google Analytics Google APIsAuth0 Inc. DoubleclickCopilot Appsflyer ChatGPT DubFengkong Auth0 Inc. AppsflyerMistral Braze Mistral OneTrust Meta ShuMeiIntercom DeepSeek Sift Grok OneTrust Singular SprigSift Science Gemini Stape DeepSeek Singular StripeStripe TikTok Twitter Ads Meta AI Twilio Segment Meta AI Twitter Analytics Wingify (a) Web client. (b) Android client.

Figure 3: Bipartite graph representing the connections to ATSes for web (left) and Android (right) clients. Legend: Interaction occurring when the user accepts non-essential cookies or ToS (Android only); Interaction occurring both before and after cookie rejection (web).

has a particularly broad footprint, with products spanning advertising and analytics (e.g., Google Ads and Google Tag Manager), search, authentication (e.g., accounts.google.com), platform APIs (e.g., subdomains under googleapis.com), and mobile-specific telemetry services such as Firebase.

Our analysis also reveals the presence of lesser-known Chinese tracking services associated with DeepSeek, including Fengkong Cloud (mobile, fp-it.fengkongcloud.com) and ShuMei (web, fp- -it-acc.portal101.cn). These services are rarely documented in the academic literature and are absent from popular tracker databases such as WhoTracks.Me. However, public reports associate them with device-fingerprinting and risk-scoring technologies [41].

In-app Browsing. 71% of all distinct endpoints contacted by mobile clients originate from WebViews (in-app browsing) rather than from native bytecode.4 Reliance on WebViews varies substantially across apps, ranging from 98% of endpoints in Grok to 91% in Perplexity and 71% in ChatGPT. In contrast, most traffic from Copilot, Mistral, and Meta originates from native code. While WebViews facilitate the integration of web content into native apps and ease development, they also enable web-based tracking methods within mobile applications [64]. In Grok, for example, we observe content from Google Ads, Google Tag Manager, TikTok Analytics, X/Twitter Analytics, and Meta Pixel loaded within WebViews.

5.2 Impact of Consent and Subscription Tiers

Consent decisions influence the presence of trackers in web-based clients. In contrast, mobile apps require users to accept the platform’s Terms of Service (ToS) and privacy policy as a prerequisite for use. We therefore evaluate the three consent scenarios described in §4.4 only on web clients: (i) ignoring the consent banner ( ignore), (ii) rejecting non-essential cookies ( reject all), and (iii) accepting all cookies ( accept all).

Our results show that consent mechanisms vary substantially across providers. Some services allow users to continue interacting without making an explicit choice (e.g., Perplexity, Claude, and Grok), whereas others require a consent choice before accessing the platform (e.g., Gemini and Meta AI). Mistral requires users to fully accept its ToS and Privacy Policy before interacting with the service, as shown in Figure 4. Therefore, all the connections are labeled as accept all. Appendix C provides examples of observed consent forms.

Even when users reject non-essential cookies (i.e., the reject all scenario), we still observe connections to third parties. For example, Perplexity, DeepSeek, Gemini, Copilot, ChatGPT, and Claude connect to Google Ads under the reject all configuration, as shown by the dashed lines in Figure 3.

Accepting all non-essential cookies activates additional third parties in Claude, Perplexity, and Grok. These correspond primarily to well-known ATSes including Meta, TikTok, Twitter Ads, DoubleClick, and AppsFlyer. These results show that some providers conditionally activate additional advertising and tracking infrastructures following user consent, as represented by the solid lines in Figure 3.

Ignoring the consent banner does not result in connections to third-party domains beyond those observed when explicitly rejecting non-essential cookies.

Subscription tiers have a comparatively limited impact on the set of third parties contacted. Free and premium accounts exhibit nearly identical tracking infrastructures across services. One exception is Claude’s mobile client, where we observe Intercom and Sentry in the free tier but not in the premium tier. However, such differences may arise from dynamically activated code paths that vary across experimental runs.

4We attribute each outbound flow to the Android UID that generated it and classify it as originating from the assistant’s native package, Custom Tabs, or a WebView process.


Table 3: Conversation artifact leakage among providers and ATSes across web and Android clients. The dissemination of shared URLs, prompts, and screenshots resulting from conversation-sharing actions is studied in §6.4.

Artifact Leaked# of Providers Web# of Providers Android# of ATSes Web# of ATSes Android
Conversation URL5090
Conversation ID2211
User Prompt1121
Conversation Title3090
Conversation Screenshot1010
Shared Conversation URL5090
Shared Conversation ID0101

Conversation ID and Shared Conversation ID counts cover only cases where the identifier leaked without the full URL.

Provisioned Tracking Surface on the Web. Observed network traffic captures only the subset of third parties activated during a particular dynamic test. However, for web-based clients, the Content-Security-Policy (CSP) response headers emitted by each service (e.g., script-src, connect-src, img-src, and frame-src) enumerate other external origins that a page is authorized to contact or load resources from, thereby revealing potential relationships that may remain dormant or untriggered during a test, for example due to regional differences or execution contexts. The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.- com) and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.

6 Privacy Analysis

Having characterized the third-party ecosystem in §5, we now study the dissemination of sensitive information from the web and mobile clients to third parties.

6.1 Dissemination of Conversation Artifacts

Unlike traditional web tracking, conversational AI platforms may leak artifacts that encode the subject matter of user conversations. permalinks, prompts, conversation titles and screenshots. Table 3 summarizes the number of third-party leaks per artifact. Overall, 6 web and 3 mobile conversational AI clients leak artifacts to 11 and 2 ATSes, respectively during regular user interactions. This includes services like Google Ads, TikTokinteractions, and Twitter Analytics as detailed in Table 4.

Conversation URLs. Conversational AI services generate persistent conversation URLs, commonly referred to as permalinks, that uniquely identify individual conversations. Conversation permalinks are stable URLs used to retrieve and manage conversation history, directly tied to a specific conversation.5 Overall, we observe that 5 web clients disclose conversation URLs or their associated global conversation identifiers to 9 ATSes. Among these, 60% (3/5) of web clients perform such disclosures by default, whereas the remaining ones only do so after users explicitly accept non-essential cookies. The web clients of ChatGPT and Claude additionally leak the globally unique conversation ID as an independent parameter to Datadog. Although less explicit than a permalink, the conversation ID allows the reconstruction of the public URL. We do not observe the conversation URL disclosure on mobile clients.

Conversation Title. Many providers automatically generate short conversation titles that summarize conversations’ content or purpose as exemplified in Table 7 in the Appendix. These titles are AI-generated summaries that concisely encode semantic information about the underlying conversation theme, purpose or intent, hence revealing sensitive interests, intentions, health concerns, financial situations, professional activities, or other personal topics discussed by the end user with the AI system. Across the evaluated services, 33.3% (3/9) of web clients leak conversation titles to 9 third parties, including Meta, TikTok, and Doubleclick. Among these, 88.9% (8/9) of them only occur when users accept non-essential cookies. This disclosure is not observed on mobile versions.

Web vs. Mobile. Many third parties appear simultaneously across both web and mobile assistants (§5), but the information they collect differs substantially. Web trackers execute within the page’s JavaScript context and can directly observe and record browser and application state, including chat URLs, page titles, conversation identifiers, and tracking cookies. In contrast, mobile clients collect conversation-related artifacts such as conversation ID, but the user-facing conversation artifacts like the titles common on the web versions are absent.

Impact of Consent and Subscription Tiers. Rejecting non-essential cookies can reduce information dissemination to third parties, as shown in Table 4. For example, rejecting non-essential cookies in Claude prevented the activation of Meta Pixel, Datadog telemetry, and server-side forwarding to eleven advertising platforms. However, across all analyzed free-tier services, third-party trackers still collect data in 44.4% (4/9) of services even when non-essential cookies are rejected. We do not observe clear differences between free and premium account tiers regarding the data collection practices of third parties.

6.2 Linking Conversations to User Identities

The disclosure of conversational artifacts becomes substantially more privacy-invasive when transmitted alongside user or device identifiers, such as hashed email addresses [25] or resettable advertising identifiers such as the Android Advertising ID (AAID) [28, 64], which enable third parties to associate those interactions with individual users or long-term cross-platform behavioral profiles [62]. We therefore analyze the extent to which conversational AI platforms transmit conversational artifacts alongside identifiers commonly used for advertising attribution, audience measurement, analytics, and cross-platform tracking by the industry.

5For example, in Grok’s conversation URL: https://grok.com/c/0b3cc700-5a98-40f0- 8e39-7311df76a700?rid=7b3e3eb2-0930-4490-b0ec-9f52fb3839c4, the path segment

and parameter uniquely identify the conversation. However, this URL is by default publicly readable for any actor knowing it, as we further develop in §6.3.


PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →