Regulation and governance for data buyers
Singapore's Computational Data Analysis Exception: Conditions for AI Training
Quick answer
Singapore's Copyright Act 2021 lets you copy works for computational data analysis, which sections 243 and 244 define to include using works to improve a program's performance, such as training an image recognition model [2]. The exception applies only if four conditions hold: copies serve CDA alone, are not supplied to others except to verify results or for collaborative research, come from a lawfully accessed source, and are not knowingly infringing [2]. It covers training, not outputs or onward licensing [1].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What sections 243 and 244 actually permit
Section 244 is a permitted use, not a license: it removes infringement liability for specific copying acts, and only while every condition is met. Section 243 frames computational data analysis around two activities: using a computer program to identify, extract and analyze information or data from a work, and using the work as an example of a type of information or data to improve the functioning of a program in relation to that type [2]. The second limb is the one AI teams rely on, because gradient updates on images, text or audio are exactly "using a work as an example" to improve a model.
The permitted use also covers copies made to prepare a work for analysis [2]. In practice that reaches format conversion, tokenization, deduplication, OCR of scanned PDFs, and resizing images into training shards, provided the resulting copies stay inside the CDA workflow. The provision has been in force since November 2021 and practitioner reviews report it unchanged since [3].
Two scoping points matter for counsel. First, the government explainer treats the exception as covering the training phase only [1]. Second, it is a defense to copyright infringement in Singapore; it does not address database rights elsewhere, trade secrets, breach of contract, or personal data under the PDPA.
The four conditions, read as engineering controls
Each condition in s244 maps to a control your pipeline must be able to evidence, and failure on any one removes the defense for the affected copies [2]. Treat them as acceptance criteria for a dataset, not as a legal afterthought.
1. Purpose limitation. The copy must be made for CDA, or to prepare the work for CDA, and must not be used for any other purpose [2]. A training shard that is later reused as a retrieval index for a RAG product, served back to users as citations, or packaged as a demo corpus has left the exception. Tag every derived artifact (raw mirror, cleaned parquet, tokenized shard, embedding store) with the purpose it was created for.
2. No onward supply. You may not supply the copy to anyone except to verify the results of the analysis or for collaborative research or study relating to it [2]. Sharing the corpus with an outside labeling vendor, a model-hosting partner, or an acquirer's diligence team needs a separate basis. Log every transfer with recipient, purpose code and the specific exception relied on.
3. Lawful access. You must have lawful access to the material you copy [2]. Circumventing a paywall or access control, or using credentials obtained outside their terms, is the classic failure. The practical question for buyers is whether whoever assembled the corpus can show how each source was accessed; see lawful access and pirated sources for the acquisition-side analysis.
4. Infringing sources. If the source copy is itself infringing, the exception survives only if you did not know and had no reason to believe it was infringing [2]. Shadow libraries, torrent mirrors and "books3"-style dumps are the scenario this condition targets. Once a source has been publicly identified as pirated, the knowledge test becomes hard to meet.
Evidence checklist for a Singapore CDA position
The checklist below turns the conditions into records a reviewer can inspect before a training run is signed off.
Illustrative example: invented to show structure; it does not describe an available dataset.
| s244 condition | Record to keep | Typical failure mode | Owner |
|---|---|---|---|
| CDA purpose only | Purpose code on each derived artifact; pipeline manifest linking shard IDs to training job IDs | Shards reused for RAG retrieval or eval leaderboards published with verbatim samples | ML platform |
| No supply to others | Transfer log: recipient, date, artifact hash, stated basis (verification or collaborative research) | Corpus copied to an annotation vendor or cloud partner bucket | Data governance |
| Lawful access | Source register: URL or channel, access method, account owner, terms version captured | Paywalled or login-gated content fetched with shared or scraped credentials | Data acquisition |
| Not knowingly infringing | Source reputation check date, takedown and blocklist screening result | Known shadow-library mirrors left in a crawl after public reporting | Legal |
| Training phase only | Separation between training store and any output, display or redistribution system | Model memorization surfaced in outputs treated as covered by s244 | Product counsel |
A training data use register is the natural home for the purpose and transfer columns, because it already tracks what each dataset may be used for.
Where the exception stops and a license is still needed
A license remains necessary whenever your use falls outside training-phase copying by you for CDA, or whenever a condition cannot be evidenced. Common triggers:
- Distribution of the dataset. Selling, sharing or open-sourcing the corpus is supply to others and outside s244 [2].
- Output and deployment uses. The explainer limits the exception to training [1]; reproduction in generated outputs, retrieval-augmented answers that quote source text, and display of training examples need their own analysis.
- Data you cannot lawfully access. Private operational records held by companies (ticket histories, engineering logs, contract repositories) are not lawfully accessible to you without the holder's permission, so the exception does not help acquire them.
- Multi-jurisdiction training. A model trained in Singapore and placed on the EU market still faces Article 53(1)(c) of the EU AI Act, which requires a copyright compliance policy that honors Article 4(3) DSM Directive reservations [5]. The UK's s29A covers only non-commercial research [4], and US analysis turns on fair use, with the Copyright Office's Part 3 report still a pre-publication version as of October 2026 [6].
- Personal data. Copyright clearance says nothing about the PDPA; see the Singapore PDPA and AI data licensing page.
Open questions counsel should track
Several points are unsettled or unverified, and a conservative position treats them as open as of October 2026. Do not assume a reading in either direction without checking the current statute text and any guidance.
- Contract override. Whether website terms of use or a database license can exclude or restrict s244 was not verified for this page. Check the Act's provisions on contractual exclusion of permitted uses before relying on CDA against terms that prohibit mining.
- Meaning of lawful access in edge cases. Publicly viewable pages behind a robots.txt disallow, rate limits or anti-bot measures are not addressed by the sources reviewed here.
- Fine-tuning versus pre-training. The definition is activity-based, so fine-tuning should fall within "improving a program's performance" [2], but retention of fine-tuning sets for later evaluation or distillation may fall outside the purpose limit.
- Case law. No Singapore judgment on AI training under s244 was identified in the sources reviewed; do not predict how a court would rule.
For a side-by-side view of how Singapore compares with other regimes, see text and data mining exceptions by country, Japan's Article 30-4 and UK commercial AI training. The Singapore country page collects the wider legal context.
How licensed operational data fits alongside s244
Licensed data complements the exception by covering what s244 cannot: records you have no lawful access to, uses beyond training, and deployments outside Singapore. A license that defines records, permitted uses, term and delivery gives you contractual evidence that travels across jurisdictions, which a Singapore-only defense does not.
When evaluating any supplier, ask for a per-dataset diligence pack covering source, rights basis, preparation (including how personal details were removed and how that was checked), and allowed use. Map those fields into the same register you use for s244 evidence so both bases sit in one audit trail; the compliance hub lists the companion records.
SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents, and finance and legal workflows, and serves AI teams wherever they are based. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Buyers can describe the data they need; a request does not guarantee a match.
Sourcing licensed training data beyond Singapore's CDA exception
If your training plan needs data that s244 cannot reach, SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees, and every release is approved by the supplying company. Start a buyer request at sourcex.si/buyers.
Sources
- Government of Singapore, "How does Singapore law treat the use of copyright works for AI training". https://isomer-user-content.by.gov.sg/61/7709dff4-3cdd-438d-b767-df5155aa943d/How does Singapore law treat the use of copyright works for AI training.pdf
- Rouse, "Artificial intelligence in Singapore: copyright infringement defence for artificial intelligence / machine learning" (2024). https://rouse.com/insights/news/2024/artificial-intelligence-in-singapore-copyright-infringement-defence-for-artificial-intelligence-machine-learning
- Reed Smith, "Text and data mining in Singapore: three years on". https://www.reedsmith.com/articles/entertainment-media-guide-to-ai-three-years-on/text-data-mining-in-singapore-three-years-on/
- UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.