What happens when your AI system knows more about your customers, employees, or users than you realize?
You might know what the model does and which AI provider you use. But do you know where personal data enters the system, where it moves next, and how long it stays there?
That is where AI compliance becomes a real DPDP question. The DPDP Act does not regulate a system simply because it uses artificial intelligence. The Act becomes relevant when the system processes digital personal data.
For businesses, this can include training data, prompts, uploaded files, RAG sources, outputs, logs, and information shared with AI vendors. Once personal data starts moving through an AI environment, the organization needs to understand what happens to it, why it is being processed, and who handles it along the way.
Note: Most of the provisions discussed in this guide, including the Act's core processing obligations and the related DPDP Rules, are scheduled to come into force on 13 May 2027. Organizations can use the transition period to prepare their controls.
When the DPDP Act Applies to AI
AI can feel like a separate technology project. From a data protection perspective, however, the technology is only one part of the discussion. You also need to look at the information moving through it.
An AI system may analyze customer support tickets, summarize employee documents, search internal knowledge bases, or help employees make decisions. If those activities involve digital personal data, the DPDP framework may apply to the processing.
Section 3 covers digital personal data processed in India, including personal data collected offline and later digitized. It can also apply to certain processing outside India when it relates to offering goods or services to Data Principals in India. Section 4 allows personal data to be processed for a lawful purpose based on consent or for certain legitimate uses listed in Section 7.
This means AI under the DPDP Act is not a separate category of compliance with its own standalone set of rules. You still need to understand the personal data involved, why you process it, and what allows you to process it.
Using AI does not answer those questions for you. Rule 13(3), however, has a direct connection with algorithmic systems. It requires a Significant Data Fiduciary, or SDF, to exercise due diligence to verify that technical measures, including algorithmic software used to process personal data, are not likely to pose a risk to the rights of Data Principals.
The Act also does not create a general right for every person to receive an explanation of an AI decision or to request human review of every automated outcome. That distinction matters because businesses can easily mix legal requirements with broader AI governance practices.
Before deciding which controls an AI system needs, it is important to understand where the personal data actually goes.
AI Data Doesn't Always Stay Where It Starts
An AI system rarely keeps data in one place. A customer may enter information into your application. Your application may send that information to an AI provider. The system may retrieve related information from an internal database. It may create embeddings, generate an output, and store logs for monitoring.
One request can create several copies of the same information. This is why AI and Machine Learning DPDPA compliance needs to begin with visibility into the data flow.
You need to understand:
- What personal data enters the AI system
- Where the data comes from
- Why did the organization collect it?
- Which model or AI service processes it
- Whether the system retrieves information from other sources
- Where prompts and outputs are stored
- Who can access the information?
- Which vendors or sub-processors receive the data
- How long does the data remain available?
Consider a simple AI assistant connected to your internal knowledge base. An employee asks a question. The assistant searches internal documents through retrieval-augmented generation, or RAG. Those documents may contain employee details, customer information, or other personal data.
The AI model may not permanently train on that information. But the system still processes it. That is why a list of AI models is not enough. You need to understand the complete flow around each model.
The same issue applies to external AI providers. A provider may act as a Data Processor when it processes personal data only on your instructions. But what happens if the provider uses your prompts or other information for its own purposes? You need to know the answer before the data enters the system.
This is also where the distinction between a Data Fiduciary and a Data Processor becomes important. AI environments can involve several parties, and each party may have a different role in the processing activity.
Once you understand where the information goes, the next step is to understand why you are using it there in the first place.
AI Training Does Not Automatically Require Fresh Consent
A common assumption is that organizations need fresh consent every time they use personal data for AI training. The answer is more complicated than that.
The DPDP Act allows processing for a lawful purpose based on consent or certain legitimate uses. When consent forms the basis for processing, Section 6 requires it to be free, specific, informed, unconditional, and unambiguous.
Consent must also relate to a specified purpose. Now consider this situation. Your organization collects customer support conversations to resolve customer issues. Later, your product team wants to use those conversations to train or improve an AI system. Simply saying, "The data belongs to us," does not justify using it for a new purpose.
You need to review why the information was originally collected, the specified purpose communicated to the individual, and whether the proposed AI use falls within the lawful processing basis available for that use.
This is where AI training data consent comes into play in India.
Before personal data enters an AI training environment, ask:
- What purpose did we communicate to the individual?
- What allows us to process the data for this AI use?
If the answers are unclear, the AI project needs another look before training begins.
Key Point to Remember : Using a dataset for an AI project does not automatically create a new legal basis. Calling the project "research" also does not automatically remove the need to review the applicable conditions.
Many AI projects also involve critical data governance decisions. After establishing why the data can be used, the next step is deciding whether all of it is necessary.
Give Training Data a Purpose, a Minimum, and an Exit Plan
Training data can move through several systems before it reaches a model. Your organization may collect the data in one system, prepare it in another, share it with an AI provider, and later use it for testing or evaluation.
That is why you need to understand where the data came from, why you collected it, and how long you need to keep it. The review should also cover what happens to the data after the AI project changes or ends.
A practical process should include:
- Purpose and Processing Basis: Establish a lawful purpose and identify the basis for processing before personal data enters the training pipeline.
- Notice and Consent: When consent is the basis for processing, Rule 3 requires that the individual receive a clear, plain-language notice containing an itemised description of the personal data being processed, the purpose of processing, and the goods, services, or uses enabled by it.
- Data Minimization: For consent-based processing, use only the personal data necessary for the specified purpose.
- Data Quality: Review whether personal data is complete, accurate, and consistent when it is likely to support a decision affecting the Data Principal.
- Security and Erasure: Protect the data and define how the organization will retain and remove it.
Your organization should also maintain clear information about the dataset itself. This can include where the data came from, when the organization collected it, why it collected it, the processing basis, and any restrictions on how teams can use it.
The data can also change as it moves through the AI environment. A raw record may later become a feature, an embedding, or part of an evaluation dataset. The system may also store related information in prompts, logs, or other technical environments.
These forms of information can remain in different systems even after the original record has been deleted.
This can create a problem when your organization needs to remove personal data. Deleting a source record may not remove the same information from a vector database, prompt log, evaluation dataset, backup, or model environment.
Retention needs to cover more than the original source record. A data retention policy under the DPDP Act can help your organization define how long personal data remains in training datasets, prompts, logs, evaluation environments, and other systems connected to the AI workflow.
The policy should also explain what happens when the retention period ends or the processing purpose no longer exists. In an AI environment, removing personal data may involve more than deleting the original record.
That is why the deletion process needs to be understood before training begins. Depending on how the system uses the data, technical deletion may involve deleting records, rebuilding an index, or assessing whether retraining or another remediation step is necessary.
The same visibility is important when you work with AI providers. Vendor agreements should explain how the provider handles personal data, including whether it can reuse the data, which sub-processors receive it, and how the provider manages security, retention, and deletion.
Once your organization understands the data used for training and how it moves through the environment, it becomes easier to review the broader DPDP requirements that apply across the AI system.
AI/ML Compliance Essentials Under the DPDP Framework
At this stage, the AI environment may involve training data, prompts, outputs, vendors, RAG systems, logs, and several internal teams. You do not need to treat every component as a separate compliance project, but you do need a consistent way to review the processing.
The table below brings together the main areas organizations should consider when reviewing AI and machine learning systems.
| Control Area | What Organizations Should Check |
|---|---|
| Legal Basis | Review why the organization processes personal data for the specific AI use. Section 4 allows processing based on consent or certain legitimate uses. AI Training Legal Basis is not a separate category under the Act. |
| Notice and Purpose | When consent is used, explain what personal data the organization processes and why. Review whether the proposed AI use fits the applicable purpose. |
| Decision Data Quality | Section 8(3) becomes important when personal data is likely to be used to make a decision affecting a Data Principal. Organizations should review whether the data is complete, accurate, and consistent. |
| Security and Logs | Security controls should cover personal data in training systems, prompts, RAG sources, outputs, and logs. Organizations should also understand what information an AI provider stores and who can access it. |
| Children's Personal Data | Organizations should review the requirements for verifiable parental consent and the applicable restrictions or exemptions before using AI for tracking, behavioural monitoring, targeted advertising, or other processing involving children's personal data. |
| Cross-Border Transfers | Understand where AI providers and cloud systems process personal data. Section 16 governs cross-border transfers. AI systems do not automatically need to keep all personal data in India, and transfer violations can attract penalties of up to ₹50 crore. |
| SDF Algorithmic Duties | Significant Data Fiduciaries should review their obligations under Rule 13(3), including due diligence related to algorithmic software used to process personal data. |
The table provides a broad view of the controls that may apply across an AI environment. However, personal data can raise different questions depending on whether the organization uses it to develop the system or processes it after the system is already in use.
That distinction becomes clearer when you separate training from inference.
Training and Inference Are Not the Same Thing
An organization can process personal data while developing an AI system and continue processing different information after the system goes live. Technical teams often describe these stages as training and inference, and the distinction is useful for privacy and compliance teams as well.
AI data privacy in India can look different depending on which stage you are reviewing.
| Area | Training | Inference |
|---|---|---|
| What Happens | The organization prepares and uses data to develop or improve a model. | The deployed model processes new requests from users or connected applications. |
| Data Involved | Training datasets, features, embeddings, and evaluation data. | Prompts, uploaded files, RAG sources, outputs, logs, and telemetry. |
| What to Check | Where the dataset came from, why the organization collected it, and what processing basis applies. | What personal data enters the system, where does it go, and who can access it? |
| AI Provider | Whether the provider can access or reuse the training data. | Whether the provider stores, accesses, or uses prompts and outputs for its own purposes. |
| Data Transformation | How personal data is transformed into features, embeddings, or other forms used during development. | How information moves through retrieval, model processing, output generation, and connected systems. |
| Retention | How long will the dataset and related information remain available? | How long prompts, outputs, logs, and other inference data remain available. |
| Data Removal | Whether deleting the original record also removes information from embeddings, indexes, or evaluation environments. | Where prompts and related information remain after processing and whether other systems contain copies. |
A Practical DPDP Workflow for AI/ML Teams
A policy can explain what your organization expects from employees and teams. It cannot tell them what to do when a new AI tool connects to customer data or when an existing vendor changes how it processes information.
Your organization needs a workflow that teams can follow.

Each step helps answer a different question.
- Describe the AI Use Case: Identify what the system does, who uses it, and whether it affects decisions about individuals.
- Map the Data: Include training data, prompts, RAG sources, embeddings, outputs, logs, backups, and vendor access.
- Identify the Parties Involved: Determine who acts as a Data Fiduciary, Data Processor, sub-processor, or service provider.
- Confirm the Purpose and Processing Basis: Check why the data is being used and what allows that use.
- Review the Data: Remove unnecessary information and check data quality where the system supports decisions affecting individuals.
- Set Security Controls: Limit access, protect stored data, and review how information moves between connected systems.
- Define Retention and Deletion: Create a Data Retention Policy under the DPDP Act for datasets, prompts, outputs, logs, and related records.
- Plan for Rights Requests: Make sure relevant systems can support requests involving access, correction, erasure, and withdrawal of consent.
- Review Major Changes: Reassess the system when the organization changes the model, data source, purpose, vendor, processing location, retention practices, or decision impact.
The level of review may also depend on the organization's responsibilities under the Act. This becomes particularly important for Significant Data Fiduciaries.
Rule 13(3) Applies Specifically to Significant Data Fiduciaries
Rule 13(3) is important because it directly refers to algorithmic software. However, this requirement applies specifically to Significant Data Fiduciaries (SDFs).
The rule requires an SDF to exercise due diligence to verify that technical measures, including algorithmic software used to process personal data, are not likely to pose a risk to the rights of Data Principals.
This creates a direct connection between Significant Data Fiduciary AI obligations and the use of algorithmic systems. Significant Data Fiduciaries also have additional obligations under the DPDP framework.
Under Rule 13(1), they must undertake a Data Protection Impact Assessment (DPIA) and an audit once every twelve months from the date on which they are notified as an SDF or included in a notified class.
A DPIA can help the organization examine how a processing activity could affect Data Principals and identify the measures needed to manage those risks before the activity continues or expands.
However, this does not mean that every organization using AI has the same obligations related to algorithmic systems. An organization that does not qualify as an SDF may still use practices such as bias testing, explainability, and human oversight.
These practices can support stronger AI governance, but organizations should not present every good governance practice as a direct requirement of the DPDP Act.
Legal obligations and internal governance practices serve different purposes. Legal requirements define what an organization must do, while governance practices may go beyond those requirements to support responsible AI use. Treating both as the same can make it difficult for teams to understand what the law actually requires.
If You Are Not an SDF, AI Compliance Does Not Disappear
SDFs have additional obligations under the DPDP framework. These include the Data Protection Officer and independent data auditor requirements under Section 10, the formal DPIA and audit, Rule 13 reporting, due diligence related to algorithmic software, and the specific transfer restriction under Rule 13(4).
Organizations that are not SDFs do not have the same SDF-specific duties. However, the general requirements for Data Fiduciaries can still apply when those provisions commence.
Depending on the AI use and processing involved, these can include processing personal data for a lawful purpose, providing notice, obtaining consent where consent is the basis for processing, or processing under certain legitimate uses where applicable.
Other requirements can include:
- Security measures
- Data quality when personal data is likely to support a decision affecting a Data Principal
- Erasure of personal data
- Handling Data Principal rights
- Breach response
- Requirements relating to children's personal data
An organization that is not an SDF may also choose to carry out an AI impact review before introducing a new AI use or making significant changes to an existing system.
This can help teams identify potential risks and consider appropriate controls. However, a voluntary AI impact review should be treated as a governance best practice unless another law, contract, or sector-specific rule makes it mandatory.
The same approach can help teams separate statutory requirements from additional governance practices. Your organization should apply the requirements that govern its processing and use additional AI controls where they help manage risks.
Four AI Compliance Claims Organizations Should Avoid
AI moves quickly, and teams may form conclusions before fully understanding how the data is being processed.
One team may think that every AI project requires consent. Another may think that personal data cannot leave India, while another may expect every automated decision to require human review.
These assumptions can create unnecessary restrictions or cause teams to overlook the requirements that actually apply.
Here are four claims your organization should examine carefully.
Treating the DPDP Act Fully in Force
Not yet. Most of the private-sector obligations discussed in this guide are scheduled to come into force on 13 May 2027.
Organizations can use the transition period to prepare their systems, data practices, and internal processes. However, they should not treat every future-effective obligation as if it already applies.
The commencement date remains important when reviewing AI compliance under the DPDP framework.
Fresh Consent for Every AI Model
Not automatically. The organization first needs to review the purpose of the processing.
Section 4 allows personal data to be processed for a lawful purpose based on consent or certain legitimate uses. When an organization introduces a new AI use, the question is whether the proposed processing fits the applicable purpose and processing basis.
A new model or AI project does not, by itself, mean that fresh consent is always required.
Right to Explanation
No, the Act does not create a general right to an explanation of every algorithmic decision or human review of every automated outcome.
Your organization may still use explainability and human oversight as part of its AI governance, and other applicable laws may create different requirements. The DPDP Act, however, should not be described as creating a general right that it does not expressly provide.
Sensitive Data Must Stay in India
Not necessarily. The DPDP Act does not use the old sensitive personal data category as a blanket requirement for data localisation.
Organizations should instead review the cross-border transfer rules that apply to their processing. Section 16 and Rule 15 address the broader transfer framework, while Rule 13(4) creates a narrower restriction for Significant Data Fiduciaries.
The organization needs to review the requirements that apply to its processing rather than assuming that all personal data in an AI system must remain in India.
Final Thoughts
AI will continue to introduce new ways to collect, connect, and use information. The organizations that handle this well will need more than a list of approved tools or a policy that teams read once.
They will need clear ownership of data decisions and a practical understanding of how those decisions change as AI systems evolve.
This can also improve how teams work together. Product teams can understand the data implications of new features, security teams can see where information moves, and privacy and legal teams can address issues while the system is still being designed.
That creates a more sustainable way to introduce AI across the business. Strong data governance can help organizations move forward with greater clarity because teams understand the information behind the technology and the responsibilities that come with using it.
Frequently Asked Questions
1. Does the DPDP Act apply to AI and machine learning systems?
Yes, when an AI or machine learning system processes digital personal data that falls within the scope of the Act. This can include training data, prompts, RAG sources, outputs, and logs containing information about identifiable individuals.
2. Does the DPDP Act provide a general right to an explanation for automated decisions?
No, the Act does not create a general right to an explanation of an algorithm or human review for every automated decision. Organizations may still use explainability and human oversight when these measures are useful or when another applicable law requires them.
3. What does Section 8(3) mean for AI-assisted decisions?
Section 8(3) requires reasonable efforts to ensure that personal data is complete, accurate and consistent when it is likely to be used to make a decision affecting the Data Principal.
Organizations should therefore pay close attention to the quality of data used in AI-supported decision workflows.
4. What does Rule 13(3) require for algorithmic software?
Rule 13(3) requires Significant Data Fiduciaries to exercise due diligence to verify that technical measures, including algorithmic software used to process personal data, are not likely to pose a risk to the rights of Data Principals.
5. Do AI systems have to keep personal data in India?
No, the DPDP framework does not create a blanket requirement for all personal data to remain in India. Organizations should review the applicable cross-border transfer requirements and any additional restrictions that apply to their specific processing activity.
6. What extra rules apply when AI processes children's personal data?
Organizations should review the requirements for verifiable parental consent and the applicable restrictions or exemptions before introducing AI-driven tracking, behavioural monitoring, targeted advertising, personalization, or other processing involving children's personal data.
7. When should organizations review an AI system for DPDP compliance?
Organizations should review the processing before introducing a new AI use case. They should also review the system when they make meaningful changes to the model, data source, purpose, vendor, processing location, retention practices, or the way AI outputs affect decisions about individuals.




Leave a Comment