Ricerca

Your sprint board is a research record. It has just never been read as one.

New architectures, novel algorithms, AI/ML systems, scale and security work. The credit was written for exactly this, and software teams are among the most likely to underclaim it. A study maps the engineering you already did to the four-part test.

Photo by Christina @ wocintechchat.com on Unsplash

Why software development qualifies

Each new or improved release is a business component - a product, software, or technique you’re trying to make better. The work qualifies when your team faces genuine technical uncertainty about architecture, algorithms, performance, scale, or security and resolves it through a systematic process of experimentation: designing alternatives, prototyping, and testing.

Every qualifying activity must pass the IRC §41 four-part test - permitted purpose, technological in nature (here, the principles of computer science), elimination of uncertainty, and a process of experimentation. Routine configuration, content changes, and bug fixes after release don’t qualify; the genuinely uncertain engineering does.

Software firms underclaim for predictable reasons: the work feels like “just building the product,” time isn’t tracked against the four-part test, and generalist preparers skip it. Mapping each activity to the statute - with contemporaneous evidence out of the systems you already run - is exactly what turns everyday engineering into a defensible claim. See documentation & substantiation for what that evidence looks like.

Seven things that happened on your roadmap last year

Not activity categories - situations. If any of these read like a sprint you actually ran, there is very likely a qualified business component underneath it.

  • “p99 latency would not come down.”

    You suspected the query planner, then the cache, then the write path. Three indexing and caching strategies were benchmarked against production-shaped traffic before one held.

    Why it can qualify: The information you had did not establish the right method. The benchmark runs are a documented process of evaluating alternatives.

  • “The model was accurate offline and useless in production.”

    Feature pipelines, retraining cadence, and drift detection were designed and re-evaluated against held-out and live data until the deployed metric matched the offline one.

    Why it can qualify: Modeling and generalization uncertainty resolved by systematic experimentation - not by reading the documentation.

  • “Tenant isolation had to survive a security review.”

    You prototyped schema-per-tenant against row-level security, load-tested both, and measured the blast radius of a misconfigured policy before committing.

    Why it can qualify: Two candidate designs, tested against requirements. Design uncertainty is one of the three flavors §41 recognizes.

  • “There was no known-good path off the monolith.”

    Strangler-fig increments, dual writes, shadow traffic, and a rollback design that had to be proven before the first cutover.

    Why it can qualify: Capability and design uncertainty about your own system. Re-architecture is a new or improved business component, not maintenance.

  • “Two systems disagreed and the connector had to decide.”

    Idempotency keys, ordering guarantees, and conflict resolution were designed and replayed against real payloads until reconciliation stopped drifting.

    Why it can qualify: Interface uncertainty resolved through iterative testing against actual data, not a configuration exercise.

  • “Usage tripled and infrastructure cost quadrupled.”

    Storage tiering, query paths, and batching strategies were each measured under load; two were abandoned after the numbers came in.

    Why it can qualify: Improving performance is a permitted purpose. The abandoned approaches are part of the experimentation, not wasted effort.

  • “You shipped it, then rewrote it six months later.”

    The first implementation worked but could not be extended; the second was built on a different abstraction after both were prototyped.

    Why it can qualify: Qualification turns on the process, not the outcome. Work on the version you replaced can still be qualified research.

Illustrative situations, not client work. Whether any of them qualifies for you depends on your facts, your contracts, and your evidence.

Two machined aluminum flange parts resting on their own engineering drawings
Design notes and rejected alternatives. Illustrative.Photo by EnCata PD on Unsplash

What the evidence looks like when it works

A defensible software claim is not assembled from memory at filing time. It is assembled from the record your team already produced: design documents that name the alternatives, benchmark runs with dates and numbers, tickets that say what was still unknown, and the branch that got abandoned.

We pull that record out of the systems you already run, tie each business component to the four-part test, and keep a line from every dollar back to the source row it came from.

How substantiation is assembled

The software work that commonly qualifies

Representative activities we see meet the four-part test. The label matters less than the underlying technical uncertainty and experimentation.

New features & architectures

Designing and building new product capabilities or re-platforming when the right technical approach isn’t known up front.

Novel algorithms & data structures

Developing algorithms, data models, or indexing and caching strategies to meet functional or performance goals.

AI/ML model development

Building, training, fine-tuning, and evaluating models, pipelines, and inference systems to resolve modeling uncertainty.

Performance & scalability engineering

Engineering for throughput, latency, concurrency, and scale when existing patterns won’t meet the target.

Security engineering

Designing authentication, authorization, encryption, and isolation to resolve technical security challenges.

DevOps & infrastructure-as-code

Developing automation, CI/CD, and infrastructure-as-code where the implementation requires experimentation.

Complex integrations

Building non-trivial third-party, API, and system integrations that require resolving interface and data uncertainty.

Prototypes & POCs

Experimental prototypes and proofs-of-concept that evaluate alternatives before committing to a design.

Platform re-architecture

Re-architecting systems for reliability, maintainability, or scale when the path forward is genuinely uncertain.

Typical QRE categories for software

What spending counts toward the credit - tailored to how software teams actually spend.

Typical QRE categories and their statutory basis
Expense category What goes into the base
Technical wages§41(b)(2)(A)-(B)W-2 wages for engineering, QA, DevOps, and technical product staff for time spent on qualified development, supervision, and direct support.
Cloud & compute§41(b)(2)(A)(iii)Amounts paid to rent cloud and compute for development, test, staging, and model-training environments used in qualified research.
Contract development (65%)§41(b)(3)65% of amounts paid to U.S. third-party developers and agencies for qualified development performed on your behalf.
Supplies§41(b)(2)(C)Limited for software teams - tangible property consumed in research, not depreciable equipment or overhead.
General and illustrative. Only qualified research performed in the United States, Puerto Rico, or a U.S. possession is eligible, and contract research enters the base at 65% of the amount paid under §41(b)(3).

What the base usually looks like

Illustrative

A directional shape for a software company, not a benchmark. Wages carry the claim; the categories below them are the ones most often left out entirely.

Technical wages - Engineering, QA, DevOps, and technical product time on qualified work.
82%
Cloud & compute - Rented dev, test, staging, and model-training capacity.
10%
U.S. contract development - Agencies and contractors, in the base at 65% of amounts paid.
6%
Supplies - Rarely material for a software team.
2%

Where the line sits

Production hosting for a released product is not research compute. Neither is the seat cost of the tools your team uses - software licenses and SaaS subscriptions are not a §41 expense category at all.

The defensible version of a cloud QRE is an allocation you can point at: tagged accounts, separate development and training projects, or environment-level billing. If nobody can tell which spend was research, an examiner will make the same observation.

Full QRE rules, category by category

Exclusions to watch

What we screen out before anything enters a base

An aggressive software claim usually fails on one of five things. Knowing which one applies to you is worth more than another list of qualifying activities.

§41(d)(4)(H)

Work a customer paid you to build

The trap for agencies, contract-development shops, and anyone doing heavy custom implementation. If the contract pays you regardless of technical success and the customer takes the rights, the research is generally funded - and funded research is out. Fixed-price work where you eat the overruns and keep the IP is a different answer.

§41(d)(4)(E)

Software built for your own back office

Software developed primarily for internal general and administrative functions faces an additional high threshold of innovation. Software you sell, lease, license, or otherwise market - and software that lets third parties initiate transactions or interact with your systems - is generally outside that internal-use category.

§41(d)(4)(A), (D)

Life after the release ships

Bug fixes, configuration, content updates, routine maintenance, and A/B tests run to answer a marketing question are not qualified research. The line is commercial production of that business component - the next genuinely uncertain improvement starts a new one.

§41(d)(3)(B)

Restyling with no technical problem underneath

A visual refresh, a brand-driven redesign, or copy and layout changes are style and taste factors. If the redesign forced a real rendering, accessibility, or performance problem to be solved, that engineering is a separate question - and it is the engineering you claim, not the palette.

§41(d)(4)(F)

Anything built outside the United States

Offshore and nearshore development is excluded from the credit no matter who employs the engineers, where the invoice is paid, or how the work is supervised. This is the single most common silent overstatement in software claims.

A credit you can use before you owe income tax

§41(h) lets a qualified small business elect to apply up to $500,000 of its research credit per year against payroll taxes instead of income tax - the employer share of Social Security tax first, and the Medicare share above that. For a venture-funded software company with a large engineering payroll and no taxable income, that is the difference between a carryforward and cash.

The definition is where companies get caught: it turns on gross receipts under $5 million for the credit year and on not having had gross receipts before the five-year window ending in that year. A company that took its first dollar of revenue six years ago is out, however small it still is. The election is made on a timely-filed return, claimed on Form 8974 with your quarterly employment tax return, and cannot be made for more than five tax years.

Rough QSB screen

  • Gross receipts under $5M in the credit year
  • No gross receipts before the five-year window ending in that year
  • A real U.S. payroll to offset
  • Election made on a timely-filed return, not after the fact
  • Five tax years is the maximum, ever

Summary only - the statutory definition and the aggregation rules decide it. We test them explicitly.

What a software study can look like

A hypothetical scenario to show how the pieces fit together. It is not a quote, projection, or promise of results.

~40-person SaaS company
Illustrative
Engineering payroll
$3.2M
Share on qualified development
~60%
Cloud & compute (dev/test/training)
$300K
Estimated QRE
~$2.2M
Illustrative federal credit
≈ $130K-$220K

Plus the full §174A first-year deduction on domestic R&E, including software development.

Illustrative only. Figures are hypothetical and rounded; the federal credit commonly works out to roughly 6-10% of QRE depending on method, filing history, and the §280C election. Your result depends entirely on your facts. This is not a quote or a guarantee.

Don’t forget §174A

Software development is fully deductible again

IRC §174A restores immediate, full expensing of domestic research & experimental costs for tax years beginning after December 31, 2024 - and domestic software development is explicitly included. Captured alongside the §41 credit, you get the deduction and the credit, correctly.

SaaS & software - frequently asked questions

Does building features for our own product count?
Generally yes, when the work meets the four-part test. Customer-facing SaaS sold or licensed to others is generally not treated as “internal-use software,” so it isn’t subject to the higher internal-use threshold - what matters is whether you were resolving genuine technical uncertainty through a process of experimentation. See the four-part test for how each element is applied.
Does AI/ML development qualify?
Often. Designing, training, fine-tuning, and evaluating models and the surrounding data and inference pipelines typically involves real modeling, performance, and capability uncertainty resolved through systematic experimentation - squarely within the four-part test. Prompt engineering against a hosted model, with no development of the underlying system, is a much weaker fact pattern.
Do cloud and compute costs count?
Yes. Amounts paid to rent or lease computing - including cloud and compute used for development, testing, and model training in qualified research - are an eligible QRE category under §41(b)(2)(A)(iii). Production hosting of a released product is not research; the allocation between the two has to be defensible, which is why we ask for tagged accounts or environment-level billing. More in qualified research expenses.
Do offshore developers count?
No. Only qualified research performed in the United States, Puerto Rico, or a U.S. possession can be claimed. Contract research counts at 65% of the amount paid under §41(b)(3), and the work must be performed there; research conducted elsewhere is excluded under §41(d)(4)(F) regardless of who employs the developers.
A customer paid us to build it. Can we still claim it?
It depends on the contract, not on the invoice. Research is “funded” - and excluded under §41(d)(4)(H) - when another party pays for it and you neither bear the financial risk nor retain substantial rights in the results. Payment contingent on technical success points toward risk staying with you; a time-and-materials agreement that pays whether or not it works, plus an assignment of all IP, points the other way. We read the agreements before anything enters a base.
Does §174A really cover software development?
Yes. Domestic software development costs are explicitly treated as research & experimental expenditures eligible for immediate expensing under §174A for tax years beginning after December 31, 2024 - see our Section 174A guide. Note that §174A eligibility and §41 credit eligibility are different tests: §174A is broader, so a well-run study captures both and does not assume the two populations are identical.
We’re pre-revenue - can we still benefit?
Possibly. A qualified small business may elect to apply up to $500,000 per year of R&D credit against payroll taxes under §41(h), turning qualified development spend into near-term cash before there is any income tax to offset. Eligibility hinges on gross receipts and how long you have had them - the mechanics are in our payroll tax offset guide.
Can we claim prior years we already filed?
Often, yes, by amending. As a general rule a refund claim must be filed within three years of filing the return or two years of paying the tax, whichever is later - so more than one prior year is frequently still open. R&D refund claims also have to meet the IRS’s specific-information requirements, which the IRS relaxed in 2024; check the current guidance before relying on any particular list. We scope lookback years alongside the current year rather than treating them as an afterthought.
What does the §280C election change?
The §280C(c) reduced-credit election lets you claim a smaller credit instead of reducing your §174/§174A deduction by the full credit amount. Which is better depends on your tax rate and posture, and it is an annual decision made on a timely-filed return. We model it both ways - see calculation methods, which also covers the regular method versus the alternative simplified credit.

Next

The four-part test, applied the way an examiner applies it

Permitted purpose, technological in nature, uncertainty, experimentation - and the shrink-back rule when only part of a release qualifies.

Also relevant

Technology & hardware

If your team ships firmware or silicon alongside the service, the hardware page carries the supply and funding rules that page does not.

See what your engineering qualifies for

Tell us about your product and stack, and we’ll map your qualifying development to the four-part test - including §174A software expensing - reviewed and finalized by R&D experts and backed by Audit Protection. Contact us for pricing tailored to your study.

[email protected] We typically reply within one business day.
Get your free credit estimate

We typically reply within one business day.