Veille IA

Doctolib, research, and professional secrecy: the promise, the clause, and mental health's blind spot

| Matthieu Ferry ⇄ IA

Doctolib reuses health data for AI research by default. What the email promises the patient, what the contract promises the practitioner, and mental health's blind spot.

The Saturday email

On Saturday, July 11, 2026, early in the afternoon, Doctolib sent its users an email entitled “Doctolib commits to research to improve health”1. It announces a research laboratory2, a first project to be conducted from August 2026 with Inria, Inserm and Université Paris Cité, and a framework: reference methodology MR-004 of the CNIL, France’s data-protection authority. Then comes a sentence, never flagged as important:

“This research project will draw on your demographic and health data necessary for scientific research purposes, as well as those of the relatives attached to your Doctolib account, whether entered by yourself or by your caregivers.”

The future tense does all the work. It asks for nothing; it informs you of a fact to come. Consent is not solicited: absent an objection, every user is included by default. Refusing requires an active step — filling in an objection form for yourself, then a separate form for each relative attached to your account.

If you are a psychologist or a psychiatrist, this email concerns you twice over. As a citizen, because your data as a patient is in the corpus. And as a practitioner, because the data you enter into your practice software — your notes, your reports, sometimes your voice dictations — is one of its sources. This article will say neither “Doctolib is plundering your data” (that is false, and believing it wastes time), nor “everything is compliant, move along” (that is true, and it is not enough). It proposes to look precisely at what is at stake, starting from the documents themselves.

This article is part of our ethics dossier — see Five principles to judge AI in mental health and our op-ed on the HAS-CNIL guide and its psychotherapy blind spot.

Three-panel infographic: the email to the patient (information by default, right to object), the clause for the practitioner (specific, circumscribed authorization), and mental health's blind spot (session verbatim, third parties mentioned, re-identification through content). Closing banner: compliance does not exhaust legitimacy.
Three panels: the email to the patient (inclusion by default, objection left to them), the clause for the practitioner (specific opt-in authorization), and mental health's blind spot (verbatim, third parties, re-identification through content). Click to enlarge.
Three-panel infographic: the email to the patient (information by default, right to object), the clause for the practitioner (specific, circumscribed authorization), and mental health's blind spot (session verbatim, third parties mentioned, re-identification through content). Closing banner: compliance does not exhaust legitimacy.

Why this is our concern

What we are pursuing here, article after article, is how to integrate AI into mental health practice without conceding anything on ethics, patient safety, or the therapeutic relationship. How a dominant player treats our patients’ data, and the effect of its choices on our clinical work, are therefore fully our subject.

Among the reservations that patients and practitioners voice about the arrival of AI in healthcare, data confidentiality comes up constantly. Our patients raise it with us; we ourselves hesitate before the “augmented” note-taking tools our software vendors offer. That worry is well-founded — but it is often misdirected, and this is the first thing to set straight.

It is not “AI” that reuses health data. A statistical model wants nothing; it decides nothing. Human structures decide: a company, a data controller, a board of directors set purposes, sign contracts, tick boxes or leave them unticked. Talking about “the algorithm” seizing our data is already picking the wrong adversary — and letting those who decide off the hook. The interest of the Doctolib case lies precisely there: nothing about it is abstract. There is a dated email, a contract published online, a checkbox in a privacy center, and signatories.

And Doctolib sets the tone. As France’s leading medical appointment platform and the publisher of widely used practice-management software, the company carries enough weight for whatever it normalizes to become the sector’s de facto standard. When a player of this size chooses a mode of data reuse, it is not merely making a corporate choice: it moves the line of what seems acceptable for everyone else.

There remains one more specific reason why this subject is ours. Frameworks designed for “general” medicine almost always ignore what makes mental health particular: a practice in which the data is not a measurement, but speech. The guide published in 2026 by France’s HAS and CNIL to support the sound use of AI in care settings illustrates this well3: organized into fact sheets by deployment stage, completed by two generic sheets on governance and generative AI, it devotes — on our reading — no specific development to the issues proper to psychotherapy: not to session verbatim, not to the third parties a patient talks about, not to re-identification through the content of a narrative. This relative silence is not a fault; it is a blind spot. And a blind spot is exactly what a practitioner of speech is best placed to name.


What is perfectly lawful

It must be said plainly, on pain of losing all credibility: Doctolib’s arrangement ticks the legal boxes.

MR-004 is a CNIL reference methodology4 that governs the reuse of already-collected health data for research “not involving the human person.” It does not require consent. It requires a public-interest purpose, data minimization, an impact assessment, a data protection officer, information of the persons concerned, and a right to object. Better still: it explicitly provides that, in the case of reuse, individual information may be replaced by referral to “a specific information mechanism” — for instance a website presenting each project. Doctolib’s “research portal” is exactly that mechanism. The arrangement does not circumvent the rule; it applies it.

And that rule has good reasons to exist. Requiring explicit consent for every reuse of health data would doom epidemiology, pharmacovigilance, and research on care pathways — the research that documents, among other things, inequalities of access. Systematic individual consent produces biased corpora: those who consent are mostly those who are doing well, who read their mail, who master the language. Opt-out — inclusion by default, absent express refusal — is not a ruse; it is a collective trade-off between each person’s autonomy and everyone’s knowledge. One may find it debatable; one cannot treat it as a scandal.

Credit must likewise be given to what is solid on the security side. The data is hosted by providers certified as “Health Data Hosts” (HDS), encrypted at rest under a dual-layer model (AES-256), with the keys kept in France in a dedicated hardware module. And the professional contract explicitly recognizes that this data is “strictly covered by professional secrecy (Article 226-13 of the French Penal Code).” Doctolib does not pretend to ignore professional secrecy: it names it.

So, if everything is lawful, where is the problem? It does not lie in any illegality. It lies in the gap between two documents that nobody reads together.


What your contract promises you

The piece that media coverage has left aside — the articles published since July 8 mostly confine themselves to the “how to refuse” instructions — is not the email sent to patients. It is the contract the practitioner signs: the data protection agreement, in its “July 2026 Version”5. Its Article 4.3 organizes the reuse of data. Here are, for the part that includes health data, the decisive passages — I abridge only the long enumeration of data categories, without altering the meaning:

(ii) Reuse of data including Patients’ health data — Subject to your specific authorization, Doctolib may reuse the following categories of data: […] your voice recordings, notably voice dictation; the personal Data of Patients […], including notably health Data, Messaging data […]. No directly identifying data will be reused.

The purpose of this reuse is to conduct research and studies; to improve and develop the Services; to anonymize the personal Data listed. […]

No authorization is presumed by default. You are free to accept or refuse, and to change your choice at any time […].

Patients’ health Data will only be processed after the collection of the consent of the Patients concerned or the obtaining of the required administrative authorizations. Patients retain the final decision on the use of their health Data: only those who have accepted, via their Doctolib account, to take part in the research […] will be concerned.

Take the time to reread that paragraph, because it holds together two promises that are hard to reconcile — and that is where everything plays out.

On one side, two very reassuring sentences. “No authorization is presumed by default”: here, for health data, it is opt-in — express, never presumed agreement — and that is a good point that deserves recognition. And above all: “only those who have accepted […] will be concerned.” A practitioner who reads this while signing understands, legitimately, that their patients will have to have said yes.

On the other side, a conjunction, two lines higher: health data will be processed “after the collection of the consent of the Patients or the obtaining of the required administrative authorizations.” That little “or” opens a second path, one that does not pass through the patients’ consent. And that second path has a name: the declaration of conformity with MR-004 is, precisely, the “administrative authorization” in question. It is the path the July program takes.

The contract does not contradict itself: it articulates two lawful routes, and it was manifestly drafted that way on purpose. But it presents them side by side without saying how they fit together. The most reassuring sentence — “only those who have accepted will be concerned” — is stated as a general principle, right after the one that reserves a form of processing without consent; and the text does not specify whether that guarantee also covers the administrative-authorization route. The important point is therefore not whether Doctolib is honoring its contract: nothing indicates that it is violating it. It lies elsewhere: could the practitioner understand, when signing, that the promise of consent was conditional? That the reassuring sentence might hold for only one of the two branches — and that it is the other branch, silent about what it implies, that would be activated?

One clarification before going further, because it changes your position in the arrangement: this second path dispenses with the patient’s consent — not with your authorization. The portal’s FAQ, addressed to the patient, says it without ambiguity: “It is only when these two conditions are met — no objection on your part and authorization from your practitioner — that the data from the professional software can be processed for research purposes.” For the data entered in your software, the Article 4.3 (ii) setting thus remains a door: as long as you have not opened it, that data does not join the corpus. You are not a spectator of the arrangement; you are one of its two locks.

This is a problem of fairness of information, not of legality. And it has a counterpart on the patient’s side. The patient was sent an email: reuse by default, objection left to them. The practitioner was made to sign a promise: “only those who have accepted will be concerned.” For the same patient, these two messages sketch two different regimes — and one of the two recipients has received information that does not prepare them for what is going to happen.

One last contrast completes the picture, and it is written in black and white in the privacy policy addressed to patients6. For health data, the announced legal basis is “the explicit consent given by the user” — immediately qualified by an exception: “unless the processing is based on public interest (research and studies).” In other words, Doctolib requires your consent to improve its products from your health data, but relies on opt-out as soon as the use is called research. The level of protection drops precisely where the purpose is presented as noblest — and where the data mobilized, verbatims and dictations included, is among the most intimate. This is not illegal; MR-004 authorizes this second regime. But it is a choice, and it deserves to be seen as one.

Let us recapitulate who holds what, because this three-actor architecture carries everything that follows:

ActorWhat they holdRegime
The patientA right to object, to be exercised for themselves and then for each attached relativeOpt-out: included by default, absent express refusal
The practitionerThe authorization to reuse the data from their software (Article 4.3 (ii))Opt-in: never presumed, revocable at any time
DoctolibThe purposes, calendar and scope of the projectsMR-004 declaration of conformity — the path that dispenses with explicit consent

What mental health changes

Up to this point, the analysis would hold for a radiology practice. It changes in nature as soon as speech is involved.

First point, technical but decisive: “not directly identifying” does not mean “anonymous.” The email says that “the data used does not allow you to be directly identified”; the portal, more rigorous, speaks of “pseudonymized” data. These are not the same thing. Pseudonymized data remains personal data — which is precisely why a right to object subsists. The CNIL and the European Data Protection Board are constant on this point — the latter devoted dedicated guidelines to it in 20257: pseudonymizing is not anonymizing. Now, the data concerned includes, as the contract states, voice dictation and messaging. In imaging, the data is a signal. In mental health, the data is a narrative — with its dates, its places, its people, its events. Re-identification through the content of such a narrative is not a textbook hypothesis.

Second point, structural and rarely raised: in a session, the patient talks about other people. About their spouse, their child, their parent, a colleague. These persons are neither Doctolib users, nor recipients of the email, nor in a position to exercise a right to object they do not know they hold. The website-based information mechanism, ingenious for the person concerned, says nothing of the third party who is described but never notified. One may object that the GDPR waives the duty to inform when the effort would be disproportionate, notably in research (Article 14(5)(b)). The exemption exists. But does it really apply to the third party whose life a session verbatim recounts? The question deserves better than an automatic referral to an exception.

Third point: the email announces “our first project.” The portal, when we consulted it in mid-July, presented two. The second — “Estimating the confidence level of artificial intelligence models” — concerns generative models. Its project sheet lists as a source “the practitioner software (including the consultation assistant)” — the tool whose function is to transcribe the consultation — and counts, among the persons concerned, the users’ “minor relatives.” This project, broader in scope than the first, is published on the portal but is not mentioned in the email. Individual information, the central obligation of MR-004, was therefore provided about the more reassuring project.

Finally, the point that concerns us most directly:

The contract does not merely make the practitioner a source of data; it makes them the one who owes patients the information.

Its Article 4.1 assigns them, as data controller, the duty to “inform the Persons concerned, notably colleagues and Patients […] by making available […] an information sheet.” This is a contractual obligation, resting on their status as data controller within the meaning of the GDPR. In other words: the portal tells the patient that their practitioner must have given authorization; the contract tells the practitioner that it falls to them to inform their patients. Each refers to the other, and the actual informing, in session, is equipped nowhere.


The shift

Let us step back. What is playing out here is not a betrayal; it is a shift, in three movements.

From consent, we move to defaults.

From deliberation, to settings.

And from the secret entrusted, to the secret deposited.

None of these movements is spectacular. Each happens through the design of a form, the order of two sentences in a contract, the state of a checkbox. That is precisely what makes them hard to grasp: there is no single, dated, signed decision one could contest. There is an architecture that makes one outcome probable — mass inclusion — while formally leaving everyone their freedom.

The logic can be summed up in one line: the burdens and the value part ways. To see it, recall that the GDPR distributes two roles: the data controller, who decides the purposes and bears the bulk of the obligations, and the processor, who acts on the controller’s behalf and under its instructions. For the activity of care, the data controller is you; Doctolib is only your processor. And everything happens as if the thankless chores of compliance were being pushed toward the practitioner, while the value — the corpus — stayed on Doctolib’s side. On the shifting of burdens, at least, the contract spells it out in black and white. Article 4.1 asks the practitioner not only to inform their patients, but to “ensure […] Doctolib’s compliance with the obligations laid down by the GDPR” and to “supervise the processing operations carried out by Doctolib.”

Supervising a provider of this size would presuppose a power of audit; yet Article 12 of the same contract hems that right in to the point of making it barely practicable: at most one audit per year, at the practitioner’s expense, with thirty days’ notice, under an agreement subject to Doctolib’s prior approval, and no intrusion testing without Doctolib’s prior written consent. Surveillance is entrusted to the party that lacks the means to exercise it. Meanwhile, for research reuse, the roles reverse: it is Doctolib that takes back the controller’s seat — and decides.

Two clauses of the general contract extend this asymmetry, and deserve to be flagged with caution, because their legal force remains to be adjudicated. The first has the practitioner declare that they had “an effective opportunity to negotiate” the contract and that they “waive any right to contest the validity of these terms on the ground of a significant imbalance.” Yet one does not waive in advance a protection of public order, and a standard-form contract offered without clause-by-clause negotiation could fall under the French regime of contracts of adhesion (Article 1171 of the Civil Code), which deems unwritten any non-negotiable clause creating a significant imbalance. The validity of such a waiver is, to say the least, debatable. These are questions for a lawyer, not for an article — but they illuminate the balance of power within which the practitioner’s famous “authorization” is given.

What is missing, at bottom, fits in one sentence: nothing in the published documents describes a watertight, audited separation between the research corpus and product development. The same subparagraph of the contract, moreover, files the two purposes together — “to conduct research and studies; to improve and develop the Services.” This is not proof of misuse; it is the absence of a guarantee. And on a corpus made of patients’ words, the absence of that guarantee is not a detail.


Is not objecting the same as agreeing?

There remains a question the arrangement sidesteps: to what, exactly, does the patient consent by keeping silent? The portal describes not a protocol but categories — “health data, lifestyle data, demographic data.” Neither the list of variables actually mobilized, nor the method, nor the impact assessment is accessible. What people are asked not to object to is a research intention more than a defined protocol.

The detailed record should nonetheless exist somewhere. MR-004 requires each project to be registered in the public directory of the Health Data Hub8, whose entries specify the exact categories of data, the sensitive variables mobilized, the objectives, the recipients. We queried that directory on July 15, 2026: of the 14,491 projects then referenced, the name “Doctolib” appears only three times, each time incidentally, never as the party responsible for a research processing operation. The project announced for August does not yet appear there in identifiable form. The most documented version — the one that would say for what, precisely — is thus not yet online. Only the portal’s lighter version remains.

And the silence commits to more than one project. The objection form is worded in general terms — objecting “to the reuse of [one’s] data […] for research purposes” — and the email already announces that this first project “paves the way for further work.” The patient who does nothing is included in the ongoing study; and since further work is announced, it will fall to them, for each project to come, to spot the information and object anew. In law, this is not a blank check: each project must be declared and the information renewed before it is implemented. In practice, the permanent vigilance thus demanded of the patient means that the absence of objection ends up resembling one.


What would make this arrangement legitimate

The question to put to Doctolib is therefore not “do you have the right?” — the answer is probably yes. It is: under what conditions would this program be legitimate? And that question calls not for a legal answer, but for a professional one. Seven conditions, which no text imposes but which the profession is entitled to demand:

1

The impact assessment made public

And not merely conducted.

2

The projects registered in the Health Data Hub’s public directory

As MR-004 provides. As of this article’s date, we have not found them there; the obligation bites before implementation, planned for August — this will be a test anyone can verify.

3

A watertight, audited separation

Between the research corpus and product training.

4

Explicit consent — an opt-in — for transcribed consultation content

Whose nature bears no comparison with an appointment history.

5

Information carried by the practitioner

With the means to carry it, rather than by an email campaign.

6

A specific regime for disciplines under reinforced secrecy

Mental health first among them.

7

Protocol-level information, project by project, and a real delay before launch

So that not objecting becomes an informed choice again, rather than an endured default.

None of these seven conditions is required by law. All seven can be demanded by the profession. That is exactly the space in which a clinician has something to say that neither a lawyer nor an engineer will say in their place.

In the meantime, there are concrete steps:

If you are a subscriber: check your setting

Open the privacy center of your professional workspace and look at the state of the data-reuse setting: as long as no one knows precisely what it covers, there is no reason to grant it — it is one of the arrangement’s two locks, and you are the one holding it.

Inform your active caseload

Nothing forbids it, and professional ethics recommends it.

If you are yourself a patient: object

For yourself and your relatives, via the portal’s dedicated form.

Take the question where it must be settled

Your learned societies, your professional board. The pouring of what your patients confide to you under secrecy into a private company’s research corpus should not be decided in a software back office.

Compliance does not exhaust legitimacy. In mental health, where the data is speech and where secrecy is the very condition of that speech, the gap between the two is at its widest. It is that gap, and not any illegality, that calls for our vigilance.


Note on method. This article relies exclusively on public sources: Doctolib’s documents — the email of July 11, 2026, the research portal, the Patients privacy policy (April 2026) and the professional contracts “July 2026 Version” (data protection agreement and general terms) — consulted in mid-July 2026; the applicable texts and reference frameworks (GDPR, reference methodology MR-004 — CNIL deliberation no. 2018-155, HAS-CNIL 2026 guide, EDPB guidelines on pseudonymization); and the public directory of the Health Data Hub, queried on July 15, 2026 through its official export (full-text search, case- and accent-insensitive). All quotations from Doctolib’s documents are our translations from the French originals. The legal analyses (significant imbalance, controller/processor articulation) are presented as questions to be adjudicated, not as settled conclusions. One element could not be observed directly: the default state of the reuse setting in the professional workspace — the article therefore makes no claim about it. All documents analyzed were archived and electronically timestamped in mid-July 2026; the timestamp certificates are kept by the author.


Références

  1. Doctolib. (2026, July 11). Doctolib s’engage dans la recherche pour améliorer la santé [email sent to users].

  2. Doctolib. (2026). Laboratoire de recherche en IA [AI research laboratory]. https://about.doctolib.fr/laboratoire-recherche-ia/ — and the research portal presenting the projects: https://about.doctolib.fr/portail-de-recherche/ (consulted in mid-July 2026).

  3. Haute Autorité de Santé & CNIL. (2026). Accompagner le bon usage des systèmes d’intelligence artificielle en contexte de soins [Supporting the sound use of artificial intelligence systems in care settings]. https://www.cnil.fr/sites/default/files/2026-03/guide_has_cnil_recommandations_ia.pdf

  4. CNIL. Reference methodology MR-004 (deliberation no. 2018-155 of May 3, 2018). https://www.cnil.fr/fr/declaration/methodologie-de-reference-04-recherches-nimpliquant-pas-la-personne-humaine-etudes-et-evaluations-dans-le-domaine-de-la-sante

  5. Doctolib. (2026). Accord sur la protection des données à caractère personnel — Doctolib Pro France [Personal data protection agreement], “July 2026 Version”. https://info.doctolib.fr/dpa/ (consulted in mid-July 2026).

  6. Doctolib. (2026, April). Politique de protection des données à caractère personnel — Patients [Personal data protection policy — Patients].

  7. European Data Protection Board. (2025). Guidelines 01/2025 on Pseudonymisation. https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2025/guidelines-012025-pseudonymisation_en

  8. Plateforme des données de santé (Health Data Hub), public directory of projects. https://www.health-data-hub.fr/ (official export queried on July 15, 2026).

Partager