The Four-Part Test in Plain English
IRC §41 asks four questions about your work. Here is what each one actually means, how examiners read them, and where real engineering projects pass or fail.
Almost every argument about the R&D tax credit - whether an activity counts, how much of an engineer’s salary is claimable, whether a study will survive an examination - resolves back to one place: the four-part test in IRC §41(d).
The test is short. Reading it in the statute is not pleasant. So here is what each part actually asks, in the words a working engineer or founder would use, along with the place each one tends to break.
One framing to carry through: the test is applied per business component, not per company and not per department. A business component is any product, process, computer software, technique, formula, or invention that you hold for sale, lease, or license, or use in your trade or business (§41(d)(2)(B)). “We do R&D” is not a claim. “This component, this year, met all four parts” is.
Part 1 - Permitted purpose
The statute: the research must relate to a new or improved function, performance, reliability, or quality of a business component (§41(d)(1)(B)(ii), §41(d)(3)).
In plain English: what were you trying to make better, and was “better” a technical property?
Those four words are the whole list. Speed, throughput, accuracy, uptime, yield, tolerance, latency, precision, capacity, durability - all comfortably inside. What is deliberately outside: style, taste, cosmetic appearance, and seasonal design factors, which §41(d)(3)(B) excludes by name. A redesign of your marketing site’s visual identity has a purpose; it is not a permitted purpose. Re-architecting the rendering path underneath it so pages paint in half the time is.
Two points teams get wrong here:
- Improving a process counts. §41(d)(3) covers a business component you use in your trade or business, so developing or improving the manufacturing process, the deployment pipeline, or the assay is on equal footing with improving the thing you sell. This is the single most under-claimed category in manufacturing - our manufacturing guide walks through what that looks like on a shop floor.
- New to you is enough. The statute never asks whether the improvement is new to the world or to your industry. Treas. Reg. §1.41-4(a)(3)(ii) makes this explicit: the fact that the information is already known to others does not, on its own, disqualify the research.
Part 2 - Technological in nature
The statute: the process of experimentation must fundamentally rely on principles of the physical or biological sciences, engineering, or computer science (§41(d)(1)(B)(i)).
In plain English: were you reasoning from a hard-science discipline, or from judgment, market feel, or aesthetics?
This is the easiest part for a technology or manufacturing company to satisfy and the easiest one to describe badly. “We used our expertise” is not an answer. “We modeled thermal behavior at the die and validated it against instrumented runs” is. The point is not sophistication - it is which discipline supplied the rules you were reasoning under.
Where it genuinely fails: pricing experiments driven by market response, A/B tests on copy, organizational process changes, and research in the social sciences, arts, or humanities, which §41(d)(4)(G) excludes outright.
Part 3 - Elimination of uncertainty
The statute: the activity must be undertaken to discover information intended to eliminate uncertainty concerning the development or improvement of the component (§41(d)(1)(A), via §174).
In plain English: at the start, did you know you could do it, know how you would do it, and know what the design should be? If all three were settled, there was nothing to research.
Treas. Reg. §1.41-4(a)(3) names the three flavors:
- Capability - could this be done at all, at the required specification?
- Method - we knew it was possible; we did not know how.
- Appropriate design - several routes existed and it was not knowable in advance which one would work.
That third flavor does most of the work in software and hardware. Teams routinely assume that because success was likely, there was no uncertainty. The standard is not risk of total failure - it is that the method or design was not knowable in advance from the taxpayer’s existing information. Choosing between two indexing strategies, three fixture geometries, or four model architectures because you could not tell from the outset which would meet the spec is textbook design uncertainty.
Where it genuinely fails: work where the answer was in a handbook, a vendor spec, or a reference implementation and you simply followed it. Configuration is not uncertainty. Neither is scale-out of a design you already proved.
Part 4 - Process of experimentation
The statute: substantially all of the activities must constitute elements of a process of experimentation (§41(d)(1)(C)).
In plain English: did you actually run a process - identify the uncertainty, identify alternatives, and evaluate them - or did you just build the first thing and ship it?
Treas. Reg. §1.41-4(a)(5) describes the shape: identify the uncertainty, identify one or more alternatives intended to eliminate it, and identify and conduct a process of evaluating the alternatives - modeling, simulation, or systematic trial and error. It does not require a lab, a hypothesis form, or a written protocol. A benchmark harness comparing three approaches, with results, is a process of evaluating alternatives.
The “substantially all” language matters and is usually read as an 80% threshold at the business-component level (Treas. Reg. §1.41-4(a)(6)): if 80% or more of a component’s research activities, measured on a cost or other consistently applied basis, constitute elements of a process of experimentation, the whole is treated as qualifying. Fall below it for the component and you fall back to the shrink-back rule - evaluate the test at the next most significant subset of elements (Treas. Reg. §1.41-4(b)(2)) rather than losing the claim entirely.
This is the part that fails most often in practice, and almost always for the same reason: the experimentation happened, and nobody recorded the alternatives. A commit history shows what was built. It rarely shows what else was tried and rejected - which is exactly what part 4 is asking for.
Three worked examples
Illustrative only - qualification is fact-specific.
A SaaS team rebuilding search. Permitted purpose: performance and quality of the product. Technological: information retrieval and distributed systems. Uncertainty: they did not know whether a re-ranking layer could hold p95 latency under the target at their index size. Experimentation: three candidate architectures, benchmarked against a fixed evaluation set, two abandoned on measured results. All four parts, in one sprint cycle. More on how this maps for software companies on our SaaS and software page.
A contract manufacturer qualifying a new alloy. Permitted purpose: quality and reliability of a process they use in their business. Technological: metallurgy and mechanical engineering. Uncertainty: feeds, speeds, and fixturing that would hold tolerance in the new material were not known. Experimentation: staged trials, dimensional inspection at each stage, scrap consumed proving the window. The trial material is a supply QRE; the trials are the evidence.
A clinical-stage biotech optimizing an assay. Permitted purpose: performance and reliability of a process. Technological: biochemistry. Uncertainty: whether sensitivity could be raised without losing specificity. Experimentation: parallel protocol variants, documented in the electronic lab notebook, with results driving the next round. See our pharmaceutical and biotech page for how this interacts with pre-revenue funding.
What the four-part test does not save you from
Passing all four parts makes an activity qualified research. It does not end the analysis. §41(d)(4) then removes several categories regardless of how well they test, and these are where otherwise-solid claims get cut:
- Research after commercial production - once the component is ready for commercial sale or use, the remaining work generally falls out (§41(d)(4)(A)).
- Adaptation and duplication - adapting an existing component to a particular customer’s requirement, or reproducing an existing component from a physical examination or plans (§41(d)(4)(B), (C)).
- Surveys, studies, and routine data collection - including efficiency surveys and routine quality control (§41(d)(4)(D)).
- Research conducted outside the United States (§41(d)(4)(F)). Note this is a location rule, not a nationality rule.
- Funded research - work funded by grant, contract, or another person, where you do not retain substantial rights and do not bear the economic risk of failure (§41(d)(4)(H) and the funded-research regulations). This is the exclusion that most often destroys a claim after the fact, because it lives in a contract nobody read for tax purposes.
- Internal-use software, which faces an additional high threshold of innovation test - innovative, significant economic risk, and not commercially available - unless it falls within an exception, such as software developed for use in an activity that itself constitutes qualified research or in a production process (§41(d)(4)(E) and the §1.41-4 regulations). The rules here are intricate; treat internal tooling as a question to answer, not an assumption.
From “it qualifies” to “we can prove it”
The four-part test is a documentation test as much as a legal one. Everything above is a question about facts that existed at a point in time, and the facts you can still show two years later are the ones you wrote down while working.
Concretely, the evidence that carries a component through the test is: what you were trying to improve (part 1), the discipline you reasoned in (part 2), what was unknown at the outset (part 3), and the alternatives you evaluated and how (part 4). If your tickets, design docs, benchmarks, test reports, and engineering change orders answer those four, you have a study. If they only show what shipped, you have a narrative problem.
Two next steps depending on where you are:
- Want the statutory detail with citations? Our four-part test deep dive and the qualified research expenses page cover the mechanics, and the R&D tax credit guide is the hub for everything else.
- Want to know what an examiner will ask for? Read what examiners actually ask for next.
Sources
- IRC §41 - Credit for increasing research activities (U.S. House, Office of the Law Revision Counsel)
- Treas. Reg. §1.41-4 - Qualified research for expenditures paid or incurred in taxable years ending on or after December 31, 2003 (Cornell LII)
- IRS - Research credit
- IRS - About Form 6765, Credit for Increasing Research Activities