The Legal Status of Synthetic Data in India’s AI Economy
Introduction : Artificial intelligence systems increasingly depend upon large and diverse datasets for training, testing and validation. At the same time, the datasets most useful for AI development may contain personal information, commercially sensitive information, confidential business material or copyright-protected works. Synthetic data has emerged as a potential mechanism for reducing the need to provide direct access to such real-world records by generating artificial data that retains selected characteristics of an underlying dataset.
The technology is also becoming relevant to India’s responsible-AI policy landscape. Under the Safe & Trusted AI pillar of the IndiaAI Mission, the Government identified Synthetic Data Generation as a specific responsible-AI theme and selected a project of the Indian Institute of Technology Roorkee concerning the generation of synthetic data for mitigating bias in datasets and machine-learning pipelines.
The growing use of synthetic data, however, raises an important legal question: does artificially generated data fall outside the legal protections attached to the original information from which it was derived?
The answer cannot be expressed through a simple rule that “synthetic data is not personal data”. The legal analysis depends upon the nature of the source material, the method used to generate the synthetic dataset, the possibility of identification or inference, and the rights that may exist in the source data and the resulting output.
Synthetic data therefore requires a lifecycle-based legal analysis:
Source data → processing and generation → synthetic output → subsequent use, disclosure and commercialisation.
Each stage can engage a different body of law.
Understanding Synthetic Data
Synthetic data is artificially generated data intended to reproduce some of the characteristics of an underlying dataset or population. NIST defines synthetic-data generation as a process in which seed data are used to create artificial data possessing some of the statistical characteristics of the seed data.
Synthetic data must be distinguished from simply modifying an existing dataset. Removing names, replacing identifiers or altering selected fields does not necessarily result in genuinely synthetic data. Such techniques may instead constitute forms of de-identification, pseudonymisation or transformation of real records.
NIST treats synthetic-data generation as one of several approaches that can be used in data-sharing and de-identification strategies. It also emphasises that organisations should assess the risks associated with releasing the resulting dataset and, where appropriate, conduct re-identification studies.
The distinction is therefore important:
Genuinely generated synthetic data consists of newly generated observations produced through a model or other generation mechanism.
Anonymised data begins with information relating to real individuals and seeks to make those individuals no longer identifiable.
Pseudonymised data reduces direct identification but may remain linkable to an individual through additional information.
A dataset described commercially as “synthetic” may therefore require closer examination before its legal status can be determined.
Synthetic Data Is Not Automatically Anonymous
The principal misconception surrounding synthetic data is that artificial generation automatically eliminates privacy concerns.
It does not.
Where a synthetic dataset is generated using a model trained on real personal data, the model may potentially reproduce or reveal information contained in the training material. The privacy risk therefore depends not only upon whether the final records appear artificial but also upon whether information relating to real individuals can be reproduced, inferred or linked from the output. NIST’s guidance accordingly recommends evaluating disclosure and re-identification risks rather than assuming that a particular de-identification technique provides complete protection.
The OECD has similarly recognised the potential of privacy-enhancing technologies to facilitate AI development while cautioning that such technologies are not “silver bullets”. Their effectiveness must be evaluated against competing considerations such as utility, efficiency and usability.
The appropriate question is consequently not:
“Is this dataset synthetic?”
but:
“Can the generation process or resulting dataset still reveal, reproduce or permit inference concerning information about identifiable individuals?”
This distinction is particularly important under India’s emerging data-protection framework.
Synthetic Data and the Digital Personal Data Protection Act
India’s principal statutory framework for digital personal data is the Digital Personal Data Protection Act, 2023 (“DPDP Act”). The Act regulates the processing of digital personal data and recognises both the protection of individuals’ personal data and the need for lawful processing.
The DPDP Act does not create a separate statutory category called “synthetic data”. Instead, its relevance depends upon whether personal data is processed during the lifecycle of the synthetic-data system.
The distinction between source data and output data is therefore central.
Consider a healthcare company that possesses a database containing patient information. It uses that database to train a model and subsequently generates artificial patient records for testing an AI system. If the resulting records cannot reasonably be connected to any patient, the output may have substantially lower privacy risk. Nevertheless, the company has still processed personal data during the earlier stages of the generation process.
The synthetic character of the final output therefore does not retrospectively determine whether the processing of the source information was lawful.
The DPDP Act defines “personal data” by reference to data about an individual who is identifiable by or in relation to such data, while “processing” covers operations performed on digital personal data. The legal analysis must consequently consider what information was used, how it was processed and whether the resulting output continues to relate to identifiable individuals.
The Current Position Under the DPDP Framework
The legal position must also be considered in light of developments after the enactment of the DPDP Act.
The Central Government notified the Digital Personal Data Protection Rules, 2025 on 14 November 2025. The Ministry of Electronics and Information Technology has also published an enforcement timeline providing for phased commencement of different provisions of the Act.
This phased commencement is important. It would be inaccurate to treat every provision of the DPDP Act as having become operational simultaneously merely because the Act was enacted in 2023.
For synthetic-data businesses, however, the underlying compliance question remains significant: an organisation should not assume that its proposed use of personal information becomes legally unrestricted simply because the ultimate product is intended to be synthetic.
The source-data processing, purpose of processing, security arrangements and subsequent handling of the output should therefore be assessed as distinct stages.
The Source-Data Problem
The most important legal issue may arise before synthetic data is created.
Suppose an Indian technology company obtains a large customer database and intends to use it to train a generative model. The company then produces an artificial dataset containing customer-like records. The final dataset may contain no names, telephone numbers or addresses.
There are nevertheless several separate legal questions:
- Was the original customer data lawfully collected?
- Was it permissible to use that data for model training?
- Was the processing consistent with the stated purpose?
- Did the organisation have the necessary contractual or other rights?
- Does the model retain or reproduce information from the training dataset?
- Can the synthetic output be linked with other information to identify individuals?
These questions demonstrate why synthetic data should be analysed across the entire data lifecycle rather than only at the point at which the output is generated.
NIST similarly recommends establishing the purpose of de-identification, evaluating disclosure risks and selecting an appropriate data-sharing model rather than assuming that the use of a particular technical method automatically eliminates risk.
Privacy, Anonymisation and Synthetic Data
Synthetic data may nevertheless provide a meaningful privacy-enhancing benefit.
Instead of giving an AI developer direct access to a real database, an organisation may provide a synthetic dataset that preserves selected statistical relationships while reducing exposure to actual individuals’ records.
This can be particularly useful in research, software testing, model development and collaborative projects where access to the original data may be commercially or legally restricted.
However, the privacy benefit depends upon the quality of the generation process. If a synthetic dataset is too closely connected to the original records, or if the underlying model memorises unusual or sensitive observations, the privacy protection may be weaker than anticipated.
The OECD therefore treats privacy-enhancing technologies as tools that can reduce the need for additional data collection and facilitate data-sharing partnerships, while emphasising the need to understand their limitations.
Synthetic data should consequently be regarded as a privacy-enhancing technique, rather than a statutory safe harbour.
Confidentiality and Synthetic Data
Privacy is only one part of the legal analysis.
A dataset may contain information that is confidential without constituting personal data. Examples include technical know-how, customer lists, manufacturing processes, commercial strategies, financial information and proprietary research.
Indian jurisprudence recognises confidentiality as a distinct legal interest. In John Richard Brady & Ors. v. Chemical Process Equipments Pvt. Ltd., the Delhi High Court considered technical information, designs, drawings and know-how that had been disclosed under conditions of strict confidence and granted relief against misuse.
The significance of John Richard Brady for synthetic data is not that the case establishes a specific rule concerning AI-generated datasets. It does not. Its relevance lies in demonstrating that confidential information may receive protection independently of the question whether copyright subsists in every element of that information.
Thus, where a business gives confidential information to a synthetic-data provider, the contractual and confidentiality relationship may remain relevant even if the resulting dataset is artificial.
Customer Databases and Commercial Confidentiality
The commercial value of databases provides another important consideration.
In Burlington Home Shopping Pvt. Ltd. v. Rajnish Chibber, the Delhi High Court dealt with a customer database maintained by a mail-order business and recognised the commercial significance of the compilation of customer information in the context of claims concerning copyright and confidentiality.
The case does not establish that every customer database is automatically protected by copyright or confidentiality. Its relevance is instead that the commercial value of a database and the circumstances in which information is compiled and used may have legal significance.
This becomes particularly relevant where an organisation creates synthetic data from a commercially valuable database. Even if the synthetic output does not contain the original customer records, contractual restrictions, confidentiality obligations and rights in the source database may still affect whether and how the transformation can be undertaken.
Copyright Protection for Synthetic Data
Synthetic data also raises an important intellectual-property question: can a synthetic dataset itself be protected by copyright?
Section 2(o) of the Copyright Act, 1957 expressly includes “tables and compilations including computer databases” within the definition of a literary work. The Act also provides that, in relation to a literary, dramatic, musical or artistic work that is computer-generated, the person who causes the work to be created is treated as the author.
These provisions are relevant to synthetic datasets, but they do not mean that every machine-generated dataset automatically receives copyright protection.
The Supreme Court’s decision in Eastern Book Company & Ors. v. D.B. Modak & Anr. is particularly important in this regard. The Court considered the originality required for copyright protection in compilations and rejected an approach based merely upon the expenditure of labour and capital.
The implication for synthetic data is significant. A business may spend substantial amounts on data collection, computational infrastructure, model training and validation. That investment may give the resulting dataset considerable commercial value, but investment alone should not be equated with copyright protection.
The copyright analysis must instead consider the nature of the compilation or other protected work and whether the statutory requirements for protection are satisfied.
Copyright in the Source Data and the Synthetic Output
The copyright analysis should also distinguish between the source material and the synthetic output.
A company may generate synthetic data from a database containing copyright-protected material. The fact that the final output is artificially generated does not automatically resolve whether the original material could lawfully be accessed, reproduced or processed for the relevant purpose.
This is particularly important where the source material contains literary works, photographs, software, artistic works or other copyright-protected subject matter.
The analysis should therefore consider:
- what copyright existed in the source material;
- whether the organisation had the necessary rights to use it;
- what part of the source material was processed;
- whether the generation process reproduces protected expression; and
- whether the resulting compilation or output independently satisfies the requirements for protection.
The distinction is crucial because copyright in the source and copyright in the output are legally separate questions.
Database Rights: The Indian Position
Copyright protection for a database should be distinguished from the concept of a separate sui generis database right.
Indian copyright law expressly recognises compilations and computer databases as literary works under Section 2(o) of the Copyright Act. However, India does not presently have a standalone database right equivalent to the sui generis regime established in the European Union.
The Indian position must therefore be assessed primarily through copyright, contractual rights, confidentiality and other applicable legal principles.
Eastern Book Company is important because the Supreme Court’s approach to originality means that the mere expenditure of labour and resources does not, by itself, establish copyright protection in a compilation.
Consequently, an Indian business should not assume that a substantial investment in creating or maintaining a synthetic database automatically creates an independent property right in the contents.
Comparative Perspective: The European Union
The position is different in the European Union.
Directive 96/9/EC on the legal protection of databases provides two distinct forms of protection: copyright protection for qualifying databases and a sui generis right for database makers where there has been qualitatively and/or quantitatively substantial investment in obtaining, verifying or presenting the contents of the database.
Article 7 of the Directive allows the maker of a qualifying database to prevent extraction or re-utilisation of the whole or a substantial part of its contents. The protection is distinct from copyright and applies irrespective of whether the database or its contents qualify for copyright protection.
The European Court of Justice has further clarified that investment in the creation of data itself is not necessarily the same as investment in obtaining, verifying or presenting existing materials for the purposes of the sui generis right. In British Horseracing Board Ltd v William Hill Organization Ltd., the Court explained the distinction between investment in creating data and investment in obtaining, verifying or presenting database contents.
This distinction is particularly interesting for synthetic data.
A business that creates a dataset through substantial computational investment may have invested heavily in generating the data itself. Under the EU framework, that does not automatically mean that all such investment qualifies for sui generis database protection. The precise nature of the investment remains relevant.
The United Kingdom Position
The United Kingdom similarly recognises both copyright protection and sui generis database rights.
UK Government guidance explains that copyright can protect the selection or arrangement of material in an original database, while database rights protect qualifying database contents where there has been substantial investment in obtaining, verifying or presenting those contents.
The UK model therefore demonstrates that copyright and database rights protect different interests.
For synthetic-data businesses, the distinction is useful because it demonstrates that a dataset can possess substantial commercial value without the legal analysis necessarily being confined to conventional copyright.
However, the EU and UK models should not be imported into Indian law by analogy. They are comparative examples rather than sources of an existing Indian database right.
What Does This Mean for Synthetic Datasets in India?
The comparative analysis exposes an important legal issue.
Suppose an Indian company spends substantial resources on:
- collecting source information;
- developing a synthetic-data model;
- generating millions of artificial records;
- validating statistical accuracy;
- removing privacy risks;
- maintaining the dataset; and
- updating it periodically.
The resulting dataset may be an extremely valuable commercial asset.
Yet the company should not assume that the level of investment automatically gives it a sui generis database right in India.
Its protection may instead depend upon a combination of:
copyright, where the statutory requirements are satisfied;
contract, where access and use are governed by agreements;
confidentiality, where the relevant information is disclosed under an obligation of confidence; and
technical controls, which restrict unauthorised access or extraction.
This makes contractual drafting particularly important for Indian businesses operating in the synthetic-data market.
Synthetic Data and India’s AI Economy
Synthetic data has significance beyond legal compliance. It can potentially address one of the structural challenges facing AI development: access to sufficiently large, diverse and useful datasets.
The IndiaAI Mission’s selection of a project specifically concerning synthetic-data generation for bias mitigation demonstrates that synthetic data is already being considered within India’s responsible-AI policy architecture.
Synthetic data may facilitate:
- AI model testing
- healthcare research;
- financial modelling;
- software development;
- privacy-preserving research;
- collaborative AI projects; and
- testing of systems where direct access to real personal data is undesirable.
The technology may therefore help reconcile data availability with privacy and confidentiality.
But its usefulness depends upon the quality of the generation process. Poorly generated synthetic data can reproduce biases, distort statistical relationships or create a false sense of privacy.
The legal and technical governance of synthetic data must therefore develop alongside its commercial adoption.
Contractual Safeguards for Indian Businesses
Given the absence of a dedicated Indian statutory regime for synthetic data, contracts can play an important role in allocating rights and responsibilities.
- Define the Source Data : The agreement should identify precisely what data is being provided to the synthetic-data provider and whether it contains personal, confidential or proprietary information.
- Define the Permitted Purpose : The provider should be authorised to process the source data only for specified purposes, such as generating synthetic data, testing a model or conducting agreed research.
- Restrict Secondary Use : The provider should not automatically be permitted to use the source data for unrelated model training, commercial analytics or development of competing products.
- Address Ownership :The agreement should expressly address rights in:
- the source data;
- the generation model;
- trained models;
- synthetic datasets;
- derivative datasets; and
- improvements to the generation system.
This is particularly important because copyright ownership and contractual rights do not necessarily produce the same result.
- Protect Confidential Information : Confidentiality provisions should cover not only the original dataset but also technical information, model architecture, proprietary methods and commercially sensitive outputs. The principles illustrated by John Richard Brady demonstrate the importance of the circumstances under which confidential technical information is disclosed and used. The principles illustrated by John Richard Brady demonstrate the importance of the circumstances under which confidential technical information is disclosed and used.
- Restrict Further Disclosure : Contracts should establish whether the synthetic dataset may be disclosed to affiliates, subcontractors, customers, researchers or other third parties.
- Provide Deletion and Retention Requirements : The agreement should specify when the original data must be deleted or returned and whether copies may be retained for security, auditing or model-maintenance purposes.
- Include Audit Rights : Businesses handling sensitive information should consider appropriate audit, reporting and compliance mechanisms.
Technical Safeguards
Contractual protection should be supplemented by technical controls.
Re-identification Testing
Businesses should assess whether individuals can be identified by combining the synthetic dataset with other information. NIST recommends evaluating re-identification risk as part of the de-identification process rather than assuming that a transformation has eliminated the risk.
Memorisation Testing
Generative systems should also be tested for whether they reproduce unusual or sensitive records from training data. This is especially important where the source data contains rare or highly distinctive information.
Differential Privacy
Differential privacy can provide formal privacy guarantees under specified conditions. It should, however, be distinguished from general claims that a dataset is simply “privacy-preserving”.
The OECD’s analysis of privacy-enhancing technologies emphasises that different technologies provide different forms of protection and that combining technologies may sometimes be necessary to address their individual limitations.
Access Controls
Source data should be subject to appropriate authentication, authorisation, encryption and logging controls.
The principle should be simple:
The synthetic dataset may be broadly accessible, but the source data used to create it should remain subject to appropriate restrictions.
Provenance and Documentation
Businesses should maintain records concerning:
- the source of the data;
- the lawful basis or authority for its use;
- the generation methodology;
- model version;
- privacy testing;
- re-identification testing;
- validation;
- permitted users;
- retention period; and
- contractual restrictions.
Such records can become important evidence of responsible governance if the legality or provenance of a dataset is later questioned.
Synthetic Data Is Not a Legal Safe Harbour
The emerging international approach demonstrates that synthetic data should be treated as one component of a broader privacy and AI-governance framework.
NIST expressly recommends evaluating the risks associated with different de-identification and data-sharing methods and does not treat synthetic-data generation as automatically risk-free.
The OECD similarly recognises the usefulness of privacy-enhancing technologies while emphasising that they are not complete solutions by themselves.
The same principle should guide Indian businesses.
Synthetic data can reduce:
- direct exposure to personal records;
- unnecessary data sharing;
- privacy risks in development environments; and
- barriers to collaborative research.
But it may also create:
- re-identification risks;
- model memorisation;
- inference risks;
- bias;
- uncertainty concerning intellectual-property ownership; and
- contractual disputes concerning the source and output.
The appropriate legal response is therefore not to prohibit synthetic data but to govern its entire lifecycle.
The Emerging Legal Position in India
The present Indian position can best be understood through the interaction of several legal frameworks rather than through a single statutory rule.
First, data protection law is relevant where personal data is processed in the creation or use of synthetic datasets.
Second, privacy principles remain relevant to the handling of information concerning individuals.
Third, copyright law may protect qualifying compilations and other protected works, while Section 2(d) addresses authorship of computer-generated works.
Fourth, confidentiality and contract law may protect commercially sensitive source information even where copyright protection is uncertain.
Fifth, database interests may arise through copyright and contractual protection, although India does not presently have the standalone sui generis database right found in the EU model.
The result is that synthetic data occupies a legally interesting position: it may reduce the privacy sensitivity of information without necessarily eliminating the legal rights or obligations associated with the source material.
Conclusion
Synthetic data is likely to become increasingly important to India’s AI economy because it offers a means of making data more usable while reducing direct exposure to real-world records.
Its legal status, however, cannot be determined merely by calling a dataset “synthetic”.
The more appropriate approach is to examine the entire lifecycle:
Source data → generation process → synthetic output → subsequent use.
At the source stage, questions concerning personal data, lawful processing, confidentiality and intellectual property may arise. During the generation stage, organisations must consider security, memorisation, inference and re-identification risks. At the output stage, questions of identifiability, copyright, database protection, confidentiality and contractual rights may become relevant.
Indian law currently does not provide a single statutory framework specifically governing synthetic data. Instead, the legal analysis must be constructed from existing frameworks including the DPDP Act, 2023, the Copyright Act, 1957, contractual principles, confidentiality jurisprudence and applicable technical and governance standards.
The copyright position also requires caution. The Copyright Act expressly includes compilations and computer databases within literary works, but Eastern Book Company v. D.B. Modak demonstrates that copyright protection cannot be based merely upon expenditure of labour and capital. The comparative position in the EU and UK further demonstrates that a separate database right can exist alongside copyright, but India’s present framework does not provide an equivalent standalone sui generis database right.
For Indian businesses, this makes contractual and technical governance particularly important. Agreements should clearly define permitted uses, ownership, confidentiality, secondary use, disclosure, retention and deletion. Technical safeguards should include appropriate access controls, provenance documentation and testing for re-identification and model memorisation.
Ultimately, synthetic data should neither be treated as inherently unsafe nor as automatically anonymous and unrestricted.
The better legal position is that synthetic data is a technological method whose legal consequences depend upon the source from which it was generated, the manner in which it was generated, the characteristics of the resulting output and the rights and obligations surrounding both the source and the output.
As India’s AI ecosystem expands, the challenge will be to ensure that synthetic data performs its intended function: facilitating innovation and responsible data use without merely transferring privacy, confidentiality and intellectual-property risks from the original dataset to a new technological layer.
Author:- Rishabh Jain, in case of any queries please contact/write back to us at support@ipandlegalfilings.com or IP & Legal Filing.
