[codicts-css-switcher id=”346″]

Global Law Experts Logo
ai training data china

Training AI Models in China (2026): Practical Compliance Guide for Personal & Sensitive Data

By Global Law Experts
– posted 40 minutes ago

AI training data china compliance has moved from a theoretical concern to an operational bottleneck for any team building models that touch personal or sensitive information within the People’s Republic of China. The evolving Cybersecurity Law (CSL) framework, combined with guidance from the Cyberspace Administration of China (CAC) and the maturing enforcement of the Personal Information Protection Law (PIPL), has raised the stakes for data science teams that curate datasets without a documented legal basis. This guide is written for in-house counsel, privacy officers and AI product legal teams who need concrete workflows rather than high-level summaries.

It sets out, step by step, how to decide your lawful basis, classify sensitive personal information, run a CAC security assessment, control cross-border transfers, and document everything for an audit.

Executive summary: what matters now (PIPL + CSL)

The core message for 2026 is that using personal data to train models in China is permitted, but only within a tightly defined compliance envelope. The interaction between PIPL (which governs personal information processing) and the CSL (which governs network security and data security obligations), together with the Data Security Law (DSL), means that a single training dataset can trigger obligations under multiple regimes simultaneously. Getting ahead of the requirements before a dataset is assembled is far cheaper than retrofitting compliance after a model is trained.

Three must-do actions frame everything that follows:

  • Establish and document a lawful basis before collection. Under PIPL you must identify a valid basis (consent, contractual necessity, legal obligation or another recognised ground) and record it before personal data enters the training pipeline.
  • Classify sensitive personal information (SPI) and treat it separately. SPI carries stricter requirements, typically separate, explicit consent plus a documented necessity test and a personal information protection impact assessment (PIPIA).
  • Determine your security-assessment and cross-border position early. Whether a CAC security assessment, standard contract filing or certification is required depends on dataset sensitivity, volume and any export of data offshore for training or hosting.

Legal bases & scope: can you use personal data for AI training data china projects?

Yes, companies can lawfully use personal data to train AI models in China, but the legality turns entirely on identifying a defensible basis under PIPL and satisfying the network and data-security duties layered on top by the CSL and DSL. The most common failure mode is not the absence of any basis, but the absence of documentation proving that a basis was chosen deliberately and applied to the specific training purpose. Regulators expect the purpose of “model training” to be identified explicitly, not buried inside a generic processing notice.

Controller vs processor roles in AI training data china pipelines

PIPL distinguishes between the entity that determines the purposes and means of processing (the “personal information processor”, functionally equivalent to a controller) and entities that process on its behalf (entrusted parties, functionally equivalent to processors). This distinction matters because it dictates who owns the lawful basis, who signs the notices, and who must sign a processing agreement.

  • If you determine why and how data is used for training, you carry primary PIPL obligations: notice, lawful basis, impact assessment, records and data-subject rights handling.
  • If you engage a vendor to label, host or generate training data, you must put an entrustment agreement in place that limits the vendor to your instructions, prohibits onward use, and requires deletion or return at the end of the engagement.
  • If you receive datasets from a third party, you must verify that the upstream party had a lawful basis to share the data for training, because inherited illegality flows downstream to your model.

Lawful bases under PIPL for AI training

PIPL sets out the grounds on which personal information may be processed. In the context of pipl ai training work, the practical hierarchy is as follows:

  • Consent. The most frequently relied-upon basis. Consent must be voluntary, informed and specific. Crucially, if the training purpose was not disclosed at collection, the original consent will not cover it, you will need fresh consent or a separate basis.
  • Necessity for a contract. Available where processing is genuinely necessary to perform a contract with the data subject, or for human resources management under lawfully adopted internal policies. Training a general-purpose model is rarely “necessary” to deliver an unrelated service, so this basis is narrower than teams assume.
  • Legal obligation or statutory duty. Applies where a law or regulation compels the processing. This is uncommon for model training but relevant in regulated sectors.
  • Public interest and other statutory grounds. Available in defined circumstances such as responding to public health emergencies or, within a reasonable scope, processing personal information already lawfully disclosed by the individual; not a general-purpose route for commercial model development.

Note that PIPL does not contain a broad “legitimate interests” basis equivalent to the one found in some other jurisdictions. This is the single most important trap for teams porting a European or US compliance approach into China: you cannot simply assert a legitimate interest in improving your product. Where consent is your basis, document the consent flow, the wording shown to users, and the timestamp, this documentation is what a regulator will ask to see first.

Where the CSL applies to csl ai data obligations

The CSL applies to “network operators”, a deliberately broad category that captures almost any organisation operating information systems in China, and imposes heightened duties on operators of “critical information infrastructure” (CII). Alongside the DSL, the framework connects network-security duties with the data-security obligations that attach to large or sensitive training datasets. Key consequences for csl ai data work:

  • Network operators must maintain security measures proportionate to the sensitivity of the data they hold, including datasets staged for training, and must comply with the Multi-Level Protection Scheme (MLPS).
  • CII operators face additional localisation and assessment duties, which can bring cross-border training arrangements within the scope of a mandatory CAC security assessment.
  • The penalty structure scales with severity, meaning that the level of a sanction depends on the sensitivity and volume of data involved and the operator’s conduct.

Sensitive Personal Information (SPI), classification & safe handling

SPI is the highest-risk category in any AI training data china project, and it is where most enforcement exposure sits. PIPL applies a stricter regime to SPI, and mishandling it is treated far more seriously than mishandling ordinary personal information.

What counts as sensitive personal information china law recognises

Under PIPL, sensitive personal information is information that, once leaked or unlawfully used, could easily lead to the infringement of the dignity of a natural person or endanger personal or property safety. The category expressly includes:

  • Biometric identifiers (facial recognition data, fingerprints, voiceprints).
  • Religious beliefs and specific identity information.
  • Medical health and financial account information.
  • Individual location tracking.
  • The personal information of minors under the age of fourteen.

Because AI training corpora frequently sweep up faces, voices, health signals and location traces, teams routinely process sensitive personal information china rules cover without realising it. A first-pass classification of every data field against the SPI list should be a mandatory gate before ingestion.

When SPI is allowed for model training

PIPL permits SPI processing only where there is a specific purpose and sufficient necessity, and with strict protective measures. In practice this means:

  • Separate consent is required unless another statutory basis applies, the consent for SPI cannot be bundled with general consent, and written consent may be required where laws or regulations so provide.
  • A documented necessity test must show that the training objective cannot reasonably be met with less sensitive data.
  • A personal information protection impact assessment (PIPIA) is mandatory before SPI is processed for training, and it must be retained.
  • Sector rules may impose additional layers, for example, health and genetic data attract further restrictions.

De-identification vs anonymisation

The distinction between de-identification and anonymisation is decisive for your regulatory footprint. Under PIPL, de-identified data remains personal information if it can be re-identified by combining it with other information; only truly anonymised data, data that cannot identify a specific individual and cannot be reversed, falls outside PIPL. National standards such as GB/T 35273 (Personal Information Security Specification) provide widely referenced technical benchmarks for the de-identification and anonymisation practices that regulators expect you to follow.

  • De-identification reduces but does not eliminate PIPL obligations; you must still perform and document a re-identification risk test.
  • Anonymisation can remove PIPL obligations, but only if you can prove irreversibility, the burden of proof sits with you.
  • Documentation is the deliverable: keep the method, the test results and the residual-risk assessment on file, because an assertion of anonymisation without evidence will not survive scrutiny.

Security assessments, filings & regulator triggers for training datasets

One of the most consequential decisions in any AI training data china workflow is whether your dataset triggers a CAC security assessment or filing. Getting this wrong, either missing a required assessment or launching one you did not need, costs weeks of programme time.

When a CAC security assessment is required for datasets

The CAC Measures for the Security Assessment of Outbound Data Transfers, read together with the 2024 Provisions on Promoting and Regulating Cross-Border Data Flows, set thresholds and triggers focused on the export of data and the handling of large or sensitive volumes. Work through this decision logic for security assessment datasets:

  • Are you exporting personal data or important data offshore for training or hosting? If yes, a security assessment or an alternative export mechanism will likely be engaged.
  • Does the dataset contain sensitive personal information or exceed the personal information volume thresholds set out in the current CAC rules? Sensitivity and scale both affect which mechanism applies and whether a mandatory assessment is triggered.
  • Are you a CII operator, or are you exporting “important data”? These trigger the security assessment route.

Because the applicable thresholds are periodically revised, confirm the current figures against the CAC’s published rules before relying on any specific number.

Assessment steps & timeline

A CAC security assessment is a structured process, and preparation is where most of the work sits. The practical sequence is:

  1. Self-assessment. Complete an internal risk self-assessment of the export or processing, documenting purpose, necessity, volume, sensitivity and recipient safeguards.
  2. Preparatory documentation. Assemble the data-processing agreement, impact assessment, technical safeguards description and data-flow map.
  3. Filing. Submit the assessment application and supporting materials to the provincial CAC, which forwards eligible applications to the national CAC.
  4. Technical review and remediation. Address any deficiencies identified, encryption gaps, access-control weaknesses, or inadequate re-identification testing.
  5. Decision. Await the assessment outcome before proceeding with the export or processing at scale.

Timelines can run from several weeks to several months depending on the complexity of the dataset and the volume of remediation required. Build this into your product roadmap rather than treating it as a formality that can run in parallel with launch.

What to expect in a regulator review

Regulators focus on necessity, proportionality and the adequacy of safeguards. Expect questions on why the data must leave China (if exported), whether less sensitive data could achieve the same purpose, and how re-identification risk has been tested. A clean, well-indexed documentation pack shortens the review; gaps invite follow-up requests that extend the timeline.

Cross-border transfers for model training & hosting

Cross-border data movement is the highest-friction dimension of ai model data transfer for training. If your compute, model-hosting or labelling capacity sits outside China, you are almost certainly exporting personal data and must select a lawful transfer mechanism.

Legal mechanisms for ai model data transfer

PIPL and the CAC measures recognise three principal routes for lawful outbound transfer:

  • CAC security assessment. The most rigorous mechanism, required for important data, CII operators, and transfers exceeding the current personal information thresholds.
  • Standard contract. The CAC standard contract provides a route for lower-threshold transfers, subject to filing with the provincial CAC and an accompanying impact assessment.
  • Certification. Certification by a CAC-accredited body offers an alternative pathway in defined circumstances.

Note that certain low-volume and specified transfers may be exempt from these mechanisms under the 2024 cross-border rules; check the current exemptions before assuming a mechanism is required.

Practical workflow for model training off-shore

The most effective way to reduce transfer risk is to design the pipeline so that raw personal data never leaves China. Practical engineering patterns include:

  • Data minimisation before transfer, export only the fields genuinely required, and strip direct identifiers where the model does not need them.
  • Federated learning, keep raw data in-jurisdiction and share only model updates, which can materially reduce the volume and sensitivity of what crosses the border.
  • Split learning and secure enclaves, partition the model so that sensitive computation stays onshore.
  • In-China training with offshore inference, train where the data lives and export only the finished model weights, subject to re-identification risk review.

Transfer recordkeeping & export logs

Maintain a transfer register recording each export: the legal mechanism relied upon, the dataset, the recipient, the safeguards and the date. These logs are among the first records a regulator requests, and their absence is treated as evidence of non-compliance.

Data governance, minimisation & technical controls for ML pipelines

Strong ai data governance china practice is what turns a defensible legal position into an operational reality. Governance must cover the full data lifecycle, not just the point of collection.

Data lifecycle map for training datasets

Map and control each stage: collection, labelling, storage, testing, use in training, and deletion. For every stage, record the lawful basis, the access controls, the retention period and the responsible owner. A lifecycle map is also the backbone of your impact assessment and the evidence you present in a security assessment.

Minimisation and synthetic data strategies

Data minimisation for ai means collecting and retaining only what the training objective genuinely requires. Practical minimisation techniques include feature selection to exclude unnecessary sensitive fields, aggregation, and, where model utility permits, substituting synthetic data generated from, but not containing, real personal information. Synthetic data can substantially lower regulatory friction, but only where you can demonstrate that it carries no realistic re-identification pathway back to real individuals.

Access, role-based controls, encryption and secure enclaves

Technical controls should be proportionate to sensitivity and documented for audit:

  • Encryption in transit and at rest for all training datasets containing personal information.
  • Role-based access control so that only named engineers can access raw or de-identified personal data, with access logged.
  • Secure enclaves and isolated training environments for the most sensitive datasets.
  • Differential privacy and tokenisation applied where they preserve utility while reducing identifiability.

Model weights & re-identification risk

Trained model weights can memorise and leak training examples. Treat the risk that a model regurgitates SPI as a live compliance issue: run hold-out and membership-inference tests, and record the results as part of your governance file.

Documentation, impact assessment templates & recordkeeping requirements

Documentation is not paperwork for its own sake, under PIPL it is a legal obligation, and it is the primary evidence that protects you in an enforcement action.

Mandatory records under PIPL

PIPL requires processors to conduct and retain personal information protection impact assessments in defined circumstances, including SPI processing, cross-border transfers, automated decision-making and entrustment of processing. Records of impact assessments and processing must be retained for at least three years. Keep a processing register that identifies each dataset, its purpose, its lawful basis, its retention period and its recipients.

Impact assessment template fields for model training

A model-training impact assessment should, at minimum, capture the following fields:

  • Purpose and necessity of the training activity.
  • Categories of data, including whether SPI is present.
  • Lawful basis and, where consent is relied upon, evidence of the consent flow.
  • Data sources and upstream lawful-basis verification.
  • De-identification method and re-identification risk test results.
  • Cross-border transfer mechanism, if any.
  • Technical and organisational safeguards.
  • Residual risk assessment and mitigation.

Retention and deletion policy

Set a retention period tied to the training purpose and delete or anonymise data once that purpose is met. Document the deletion, including deletion from backups and derived datasets, because a retention breach is a common and easily proven enforcement trigger.

Enforcement risk, penalties & liability scenarios

Understanding how enforcement actually arises helps prioritise controls. The regulator’s attention tends to concentrate on a small number of recurring failures.

Typical enforcement triggers

  • Unauthorised cross-border export of personal or important data without the required assessment or mechanism.
  • SPI misuse, processing biometric, health or location data without separate consent or a documented necessity test.
  • Purpose creep, using data collected for one purpose to train a model without a fresh basis.
  • Absent documentation, inability to produce an impact assessment, processing register or transfer log on request.

Penalties, remediation & crisis response checklist

PIPL provides for significant administrative fines, orders to suspend or terminate processing, confiscation of unlawful gains, and, in serious cases, personal liability for responsible individuals and potential criminal exposure under the Criminal Law. For grave violations, PIPL authorises fines of up to RMB 50 million or up to 5% of the preceding year’s turnover, alongside possible business suspension and licence revocation. If an incident or investigation arises:

  1. Preserve evidence and freeze the affected pipeline.
  2. Assemble the impact assessment, processing register and transfer logs for the affected datasets.
  3. Assess and, where required, notify affected individuals and the regulator.
  4. Remediate the root cause and document the remediation.
  5. Engage local counsel before responding substantively to the regulator.

Compliance playbook: step-by-step checklist & roles for AI training data china projects

The following staged workflow assigns responsibilities across legal, privacy, data engineering and security functions:

  1. Prepare. Legal maps the lawful basis; privacy classifies SPI; engineering builds the data-flow map.
  2. Assess. Complete the impact assessment and determine whether a CAC security assessment or transfer mechanism is required.
  3. Remediate. Security implements encryption, access control and enclaves; engineering applies de-identification or synthetic-data substitution.
  4. Document. Populate the processing register, impact assessment and transfer logs.
  5. Monitor. Run re-identification and leakage tests, review retention, and re-assess when the purpose or dataset changes.

The decision below sits at the centre of this playbook: how you source your training data, raw personal data, de-identified data, or synthetic data, drives every downstream obligation.

Dimension Use personal data (raw) Use de-identified data Use synthetic data
Legal basis required (PIPL) Strongest constraints, often explicit consent or strong necessity plus impact assessment; SPI needs stricter approvals Still personal data if re-identifiable; needs measurable de-id standard plus impact assessment If irreversibly non-personal, PIPL may not apply; document generation and test residual re-identifiability
Need for CAC security assessment High if size, sensitivity or cross-border transfer involved Conditional on re-identifiability and export destination Lower likelihood; regulator will scrutinise synthetic methods and linkage risk
SPI exposure risk High, extra safeguards, separate consent, higher penalties Medium, mitigations reduce but do not eliminate risk Low, but must demonstrate no real-world linkage
Cross-border transfer complexity Highest, likely assessment or certification Medium, fewer transfers if hosted in China Low to moderate depending on generation source
Recordkeeping & impact assessment Full PIPL logs, impact assessment and processing records mandatory Impact assessment required; document re-id risk testing Impact assessment on generation and residual risk; model cards recommended
Engineering controls Encryption, secure enclaves, limited access De-identification, tokenisation, differential privacy Synthetic pipelines, audit logs, hold-out leakage tests
Timeline impact Long: consent, assessment, remediation (weeks–months) Medium: de-id testing and validation (weeks) Shorter if generation and testing are sound (days–weeks)
Enforcement risk Highest (fines, suspension, criminal liability in egregious cases) Medium, mitigatable with documentation Low if demonstrably non-personal

Decision framework, which approach to choose

Take a position rather than hedging. The recommended default for most commercial teams is de-identified data, escalating to synthetic data where SPI or cross-border risk is high, and reserving raw personal data for cases where nothing else will do:

  • Choose raw personal data only when you hold a lawful basis that cannot be substituted, data subjects have given the required consent for training, and you can complete any required CAC assessment and technical safeguards.
  • Choose de-identified data when you can remove direct identifiers and prove low re-identification risk through testing, and you need more model utility than synthetic data provides with lower regulatory friction. This is the pragmatic default.
  • Choose synthetic data when utility needs are modest, SPI risk is high, or cross-border training is required and you want to avoid export assessments, but only after validating that no realistic re-identification pathway exists.

A compliance checklist and impact-assessment template for AI training accompanies this guide as a working resource for your team.

Sector examples & short case studies

Connected vehicles

Connected-vehicle programmes generate continuous location traces, in-cabin video and biometric signals, much of which qualifies as SPI, and some of which may be treated as important data under automotive-data rules issued by the CAC together with the MIIT and other authorities. A manufacturer training a driver-monitoring model should keep raw video in China, de-identify or synthesise faces before any offshore processing, and expect a CAC security assessment if regulated telemetry crosses the border. Location tracking of individuals is expressly sensitive, so separate consent and a necessity test are non-negotiable.

Healthcare and clinical AI

Medical health information is SPI, and health-sector rules on human genetic resources and medical data add further restrictions on top of PIPL. A clinical-AI developer should assume that separate consent, an impact assessment and rigorous de-identification are all required, and that anonymisation claims will be tested against recognised national standards. Federated learning across hospital sites, keeping patient records in place and sharing only model updates, is often the most defensible architecture for reducing both transfer and SPI exposure.

Advertising and audience modelling

Advertising teams frequently rely on behavioural and location data whose original consent did not contemplate model training. The recurring failure here is purpose creep: repurposing profiling data for a training corpus without a fresh basis. The safer path is aggregated or synthetic audience data, supported by an impact assessment that documents the minimisation applied and the re-identification testing performed.

Conclusion & next steps

Managing AI training data china compliance in 2026 is an operational discipline, not a one-off legal sign-off. The combination of PIPL’s strict lawful-basis and SPI rules with the CSL and DSL security obligations and the CAC’s cross-border assessment triggers means that the teams who succeed are those who classify data, document their basis, test de-identification and control cross-border movement before a single model is trained. Adopt de-identified or synthetic data as your default, reserve raw personal data for cases of genuine necessity, and keep a complete impact assessment, processing register and transfer log for every dataset.

When your project involves SPI, cross-border training or a potential CAC security assessment, engage qualified local counsel early, the cost of doing so is trivial against the cost of a suspended pipeline or an enforcement action. For further reading, see the China data protection lawyers 2026 (practice overview) and the GLE lawyer directory, Data Protection lawyers in China. Supporting resources on cross-border data transfers (China), impact assessment templates for AI training, and sector checklists for connected vehicles and healthcare accompany this pillar guide.

Need Legal Advice?

This article was produced by Global Law Experts. For specialist advice on this topic, contact Maggie Meng at Beijing Global Law Office, a member of the Global Law Experts network.

Sources

  1. Personal Information Protection Law of the PRC (PIPL), National People’s Congress
  2. Cybersecurity Law of the PRC, National People’s Congress
  3. Data Security Law of the PRC, National People’s Congress
  4. Cyberspace Administration of China (CAC), Measures for the Security Assessment of Outbound Data Transfers and Provisions on Promoting and Regulating Cross-Border Data Flows
  5. Ministry of Industry and Information Technology (MIIT), sectoral data and automotive guidance
  6. OECD, AI Principles and policy guidance

FAQs

Can companies use personal data for AI training data china projects at all?
Yes. PIPL permits using personal data to train models where you have a valid lawful basis, most commonly consent, and satisfy the network and data-security duties imposed by the CSL and DSL. The purpose of “model training” must be identified explicitly and documented before collection.
PIPL requires separate consent for processing sensitive personal information unless another statutory basis applies, together with a documented necessity test and a personal information protection impact assessment. Written consent may be required where laws or regulations so provide. Bundling SPI consent into a general notice will not satisfy the standard.
A security assessment is required where you export important data, where you are a critical information infrastructure operator, or where personal information exports exceed the current volume thresholds in the CAC rules. Lower-threshold transfers may instead use a standard contract or certification, and some low-volume transfers may be exempt. Work through the current CAC Measures and cross-border provisions early and confirm the applicable thresholds before relying on any figure.
No. De-identified data remains personal information under PIPL if it can be re-identified. Only truly anonymised, irreversible data falls outside PIPL, and the burden of proving irreversibility, following recognised national-standard testing, sits with you.
Keep raw data in China and export as little as possible. Federated learning, split learning, in-China training with offshore inference, and synthetic-data substitution all reduce what crosses the border and can lower or remove the need for a security assessment.

Find the right Legal Expert for your business

The premier guide to leading legal professionals throughout the world

Specialism
Country
Practice Area
LAWYERS RECOGNIZED
0
EVALUATIONS OF LAWYERS BY THEIR PEERS
0 m+
PRACTICE AREAS
0
COUNTRIES AROUND THE WORLD
0
Lawyer Profile Page - Lead Capture
GLE-Logo-White
Lawyer Profile Page - Lead Capture

Training AI Models in China (2026): Practical Compliance Guide for Personal & Sensitive Data

Send welcome message

Custom Message