Ricerca
AISoftwareR&D Credit

R&D Tax Credit for AI and Machine Learning Development

Which AI and ML work qualifies for the R&D credit, which doesn't, and how wages, cloud GPU compute and contract research count, with the records to keep.

The Ricerca Team 11 min read

AI and machine learning development can qualify for the federal R&D credit under IRC §41 when your team resolves technical uncertainty through a systematic process of experimentation: designing or adapting model architectures, developing training methods, building ways to evaluate a system whose behavior you cannot predict, and testing alternatives against measurable targets. What usually does not qualify is using a model as-is, adjusting prompts without a designed experiment, routine data labeling, and work done after a model is in production. The qualifying costs are mostly wages of the people doing the experimentation, cloud compute rented for training and experiments, and 65% of U.S. contract research.

If you are an early-stage founder weighing the payroll tax offset and §174A expensing, our guide to R&D tax credits for Silicon Valley AI and SaaS startups covers that ground. This post is narrower: where the qualification line falls for ML work, how each cost is treated, and what records hold up.

How the four-part test reads for an ML project

The credit is applied per business component - a product, process, software, technique, formula or invention (§41(d)(2)). For an AI company that might be a model, a training pipeline, an evaluation system or a feature built on top of them. Each one has to pass all four parts of the four-part test, and in ML each part has a specific meaning:

  1. Permitted purpose. The work aims at a new or improved function, performance, reliability or quality. Accuracy, latency, cost per inference, robustness and failure rates are all measurable purposes.
  2. Technological in nature. The experimentation relies on principles of computer science, engineering, or the physical or biological sciences. Machine learning work usually rests on computer science. Research into why users behave as they do is social science, which is excluded.
  3. Elimination of uncertainty. At the start, you did not know whether you could reach the target, how to reach it, or what the right design was. Treas. Reg. §1.41-4(a)(3) is explicit that you do not have to go beyond what skilled professionals in the field already know, and you do not have to succeed. The uncertainty is judged on the information available to you.
  4. Process of experimentation. You identified the uncertainty, identified alternatives, and evaluated them through modeling, simulation or systematic trial and error (Treas. Reg. §1.41-4(a)(5)). Substantially all of the activity - 80% or more, measured on cost or another consistent basis - has to be part of that process.

The fourth part is where ML claims are won or lost. Training runs are cheap to launch and easy to repeat. An examiner will want to see that they were designed to answer a question, not just rerun until something looked good.

AI work that tends to qualify

In practice the strongest ML components share a shape: a stated target, a technical reason it was hard, and a record of alternatives that lost. Examples:

  • Architecture decisions under real uncertainty. Designing a new model structure, or significantly modifying an existing one, when you could not know in advance whether it would hit your accuracy or latency target.
  • Training methods. Developing loss functions, data mixtures, curricula, distillation or fine-tuning strategies where the right approach had to be found by controlled comparison.
  • Evaluation you had to invent. Building a way to measure a system whose outputs are not deterministic, when no existing benchmark captured the failure you cared about.
  • Systematic experiments across alternatives. Ablations, controlled comparisons of retrieval or ranking approaches, or hyperparameter searches designed to answer a specific technical question.
  • Data engineering that changes the model. Synthetic data generation or filtering methods developed and tested to fix a measured failure mode.

What usually doesn’t qualify

The boundary is just as important, and most of it comes straight from the statute:

  • Using an off-the-shelf model as-is. Calling a hosted model through its documented API, or deploying a pretrained model without technical change, rarely involves uncertainty about capability or method. The regulations add that using computers to process data “does not itself establish that qualified research has been undertaken” (Treas. Reg. §1.41-4(a)(7)).
  • Prompt tweaking without a process of experimentation. Adjusting wording until the output reads better is iteration, not an evaluative process. Tuning a chatbot’s tone to match a brand voice is closer to the style and cosmetic factors that §41(d)(3)(B) excludes than to function or performance.
  • Routine data labeling and collection. The statute excludes “routine data collection” by name (§41(d)(4)(D)(iv)). Labeling done as ongoing operations is not research, even when the labels later feed a model.
  • Work after the model ships. Research after commercial production is excluded (§41(d)(4)(A)). Scheduled retraining on fresh data with an unchanged pipeline, monitoring and routine tuning usually fall here.
  • Adapting a model to one customer’s needs. Adapting an existing component to a particular customer’s requirement is excluded (§41(d)(4)(B)), although the exclusion does not apply merely because work is intended for a specific customer (Treas. Reg. §1.41-4(c)(3)).
  • Behavioral and market testing. A/B tests aimed at conversion, user surveys and similar work are market research or social science (§41(d)(4)(D)(iii) and (G)).
  • Research performed outside the United States (§41(d)(4)(F)), including by an offshore ML team.

Which costs count as QREs

Once a component qualifies, the credit is computed on qualified research expenses. For AI teams, four categories matter.

Wages. W-2 wages of ML engineers, research scientists and data scientists for time spent performing the experimentation, first-line managers who directly supervise it, and staff who directly support it (§41(b)(2)(B)). The regulations list “a clerk for compiling research data” as direct support, so data preparation done for a specific experiment can count, while labeling done as routine operations generally does not. If 80% or more of a person’s wages are for qualified services in a year, all of that person’s wages count (Treas. Reg. §1.41-2(d)(2)).

Cloud GPU and compute. The statute allows amounts paid to another person “for the right to use computers in the conduct of qualified research” (§41(b)(2)(A)(iii)). Treas. Reg. §1.41-2(b)(4) sets three conditions: the computer must be owned and operated by someone else, located off your premises, and you must not be its primary user. Shared cloud capacity used for training runs, ablations and evaluation in qualified research often fits. Three limits apply:

  • Production inference does not count. Compute that serves customers is not used in the conduct of qualified research.
  • Dedicated hardware raises questions. Reserved or dedicated instances can raise the primary-user question, so they deserve a closer look.
  • Resale reduces the amount. The amount is reduced by anything you receive for the right to use substantially identical property, which matters if you resell capacity.

GPUs you own are depreciable property. They are neither supplies nor computer rental, so their cost is not a QRE.

Supplies. Supplies must be tangible property (§41(b)(2)(C)). Licensed datasets and software subscriptions are intangible, so they are not supply QREs. Whether per-token fees for a hosted model used in experiments count as computer rental is not settled by the regulations. Treat it as a question for your study, not an assumption.

Contract research. Generally 65% of amounts paid to another person for qualified research performed on your behalf (§41(b)(3)(A)). Under Treas. Reg. §1.41-2(e), the agreement must be in place before the work, you must have a right to the results, and you must bear the cost even if the research fails. Payments contingent on success are for a product, not research. The work must also be performed in the United States.

Building AI for clients: the funded research question

If you build models for clients, the biggest risk to your claim is in the contract. Research is excluded to the extent it is funded by a grant, a contract or another person (§41(d)(4)(H)). Two questions decide it: do you keep substantial rights in the results, and do you bear the economic risk if the work fails? A time-and-materials engagement where the client owns everything and pays regardless of outcome usually fails both. A fixed-price build where you absorb overruns and keep rights in the methods can survive. Our post on how contract clauses decide funded research goes clause by clause.

Internal AI tools and the internal-use software rules

An AI tool you build for your own back office is internal-use software if it serves general and administrative functions: finance, HR or support services, which include marketing and data processing. A resume-screening model, a finance forecasting tool or an employee knowledge assistant would fall here. These must pass the stricter high threshold of innovation test as well as the four-part test.

Customer-facing AI features are not internal-use software. Two exceptions also switch off the higher test: internal software developed for use in an activity that is itself qualified research, such as an evaluation platform your team uses in qualifying model development, and internal software for a production process that itself meets the four-part test (Treas. Reg. §1.41-4(c)(6)(ii)).

§174A: the deduction side of AI development

The credit and the deduction are separate benefits. Under IRC §174A, any amount paid or incurred in connection with developing software is treated as a research or experimental expenditure, and domestic research spend is deductible in the year paid or incurred for tax years beginning after December 31, 2024. Foreign research spend is still amortized over 15 years under §174. The §174A category is broader than the credit: routine software development can be deductible under §174A even though it earns no credit.

The two interact. Under §280C(c), your §174A deduction is reduced by the amount of the credit unless you elect the reduced credit on a timely filed original return. Our §174A vs §41 comparison lays out the interplay.

How to document ML experimentation

ML teams already produce much of the evidence a study needs. The gap is usually linking it to business components and people. Keep:

  • Experiment logs recording the hypothesis, configuration, dataset version, metric, result and the decision that followed.
  • Evaluation results with baselines and the thresholds you were trying to beat.
  • Model cards and technical reports describing design choices, rejected alternatives and known limits.
  • Training run records exported from your tracking tools, plus the related commits, pull requests and tickets.
  • Time allocation showing who worked on which component.
  • Cloud bills split by project or account, so research compute can be separated from production.

The revised Form 6765 asks for business-component-level detail in Section G. Under the current Instructions for Form 6765, Section G is required for tax years beginning after 2025, with exceptions for qualified small businesses electing the payroll credit and for smaller filers claiming on an original return. Records that already map experiments to components make that section straightforward. A Ricerca study connects payroll, GL and engineering systems (or works from uploaded exports), and R&D experts finalize every study. Your CPA or tax preparer signs and files the return.

FAQ

Does fine-tuning an open-weight model qualify?

It can. Fine-tuning qualifies when you were uncertain whether or how the model could meet a measurable target and you evaluated alternative approaches to find out. Following a documented recipe to a result you already expected is much weaker.

Do cloud GPU costs count as QREs?

Often, for compute used in qualified research. The computer must be owned and operated by the provider, located off your premises, and you must not be its primary user. Compute that serves production traffic does not count.

Does prompt engineering qualify?

Only when it is part of a systematic, evaluative process aimed at a technical uncertainty, with defined metrics and alternatives tested against them. Ad hoc wording changes and tone adjustments do not.

Is our internal AI assistant eligible?

If it serves administrative functions such as HR, finance or internal support, it is internal-use software. It can still qualify, but only if it also passes the high threshold of innovation test.

For the fundamentals across industries, see the R&D tax credit guide and our SaaS and software industry page.

Sources

R&D tax credit updates

Plain-English notes on the R&D credit, §174A and state credit changes. About twice a month.

We will email you a link to confirm.

See what your engineering year is worth

Wages, cloud and compute used in development, and qualifying contract research - computed from your actual systems, then reviewed by R&D experts before anything is issued.

[email protected] We typically reply within one business day.
Get your free R&D credit assessment

We typically reply within one business day.