Two years ago, the IRS ran about 10 active AI projects. Today it runs 126 — and six times a year, a set of machine-learning models pores over every small business and self-employed tax return filed, looking for patterns a human examiner would take weeks to spot. On February 10, 2026, the agency made this official by codifying the practice into the Internal Revenue Manual as IRM 10.24.1, its first formal policy governing how AI selects returns for audit.
If you run a business and file a Schedule C, an 1120-S, or a partnership return, this isn't background noise. It's a real shift in who gets a second look from the IRS, and why. Here's what changed, what the algorithms are actually looking for, and how to make sure your books never give them a reason to flag you.
What IRM 10.24.1 Actually Says
The new manual section does three concrete things:
- Defines what AI is authorized to do in the examination process — including, explicitly, systems that "inform or influence whether a taxpayer will be subject to audit, or what aspects of a return will be subject to audit." The IRS calls this a "presumed high-impact" use of AI, putting it in the same risk category as systems that touch civil rights or critical infrastructure.
- Requires mandatory human review before an AI-generated referral becomes an actual audit. A model can flag a return, but a person has to sign off before it turns into a notice.
- Sets documentation requirements for AI-assisted case selection, so examiners have to be able to explain, in writing, why a case was picked.
The policy also creates a governance structure: a Chief Data and Analytics Officer serves as the IRS's "Responsible AI Official," and a board called the Data and Analytics Strategic Integration Board signs off on any new high-impact AI use case before it goes live.
In plain English: the IRS isn't hiding the fact that algorithms are doing first-pass triage on your return. It's now writing down the rules for how that triage works — which also means examiners have to be able to justify it later.
How the Models Actually Pick Your Return
Two of the IRS's AI models are aimed specifically at self-employed taxpayers and small business owners, and for a simple reason: underreported self-employment income is the single largest component of the individual tax gap. These models don't look at your return in isolation — they compare it against your own filing history and against thousands of similar businesses in your industry. A few things reliably raise the score:
- Year-over-year discrepancies. A consulting business that reported $180,000 in gross receipts last year and $95,000 this year, with no obvious explanation, stands out immediately.
- Deduction ratios outside industry norms. The IRS holds statistical profiles of typical expense-to-revenue ratios by industry. If your business claims 65–70% of gross receipts as deductions when the norm for your NAICS code runs 35–45%, that gap gets flagged — even if every deduction is legitimate.
- Round numbers. A schedule showing office supplies at exactly $5,000, utilities at $3,000, and meals at $2,000 reads as estimated, not recorded. Real bookkeeping produces numbers like $4,847.13 and $2,963.82, not tidy multiples of five.
- Information mismatches. The IRS cross-references W-2s, 1099s, and K-1s it receives directly from payers. If a 1099-NEC for $42,000 doesn't show up anywhere on your return, that's an automatic, near-instant mismatch — no AI required to catch it, but AI now helps prioritize which mismatches matter most.
- Mileage and expense outliers. Claims like 40,000+ business miles a year for a single-vehicle business, or meals-and-entertainment costs above 10% of revenue, are statistical outliers that get weighted heavily.
None of this is new in kind — the IRS's Discriminant Function (DIF) scoring system has compared returns to statistical norms for decades. What's new is the scale and speed: instead of one static formula reviewed periodically, you now have models that re-run six times a year and get better at spotting patterns with every cycle.
The Part That Should Worry You More: Response Time, Not Detection
Here's the asymmetry small business owners are actually running into in 2026: the AI-assisted flagging is fast, but the human side of resolving a flag hasn't scaled with it. Automated notices go out quickly. Getting a live person on the phone, or a substantive written response to a notice, is now taking noticeably longer than it did two years ago. A flag that would have been a five-minute phone call in 2023 can now sit unresolved for months while penalties and interest accrue.
That means the real cost of a messy return isn't just audit risk — it's the time and stress of untangling a flag once it's raised, with a slower human backstop on the other end. The best defense is not giving the algorithm a reason to flag you in the first place.
Why the IRS Had to Write This Policy Down
The governance structure in IRM 10.24.1 isn't just bureaucratic housekeeping — it's a direct response to a real problem with algorithmic audit selection. Independent research and Government Accountability Office reviews have found that certain groups of taxpayers, including Black taxpayers claiming the Earned Income Tax Credit, have been audited at rates three to five times higher than other filers with similar returns. The GAO's working theory is that this isn't intentional targeting but unintentional algorithmic bias — models trained on historical audit data can end up reproducing whatever selection patterns existed in the past, even when nobody designs them to.
That's precisely why the new policy requires impact assessments, ongoing monitoring, and mandatory human sign-off before a model's referral becomes an actual audit: a machine-learning system optimized purely for "catch underreported income" has no built-in check on whose income it disproportionately flags. Small business owners should read this as a reason for cautious optimism rather than a lawsuit-in-waiting — the human-review requirement is a real backstop — but it also confirms that the AI's first-pass score isn't infallible. If you're flagged, it doesn't mean the algorithm found something real; it means a pattern-matcher found something statistically unusual, which your own records may explain perfectly well.
If You Do Get Flagged
Getting a notice doesn't mean you did anything wrong — it usually means one number, somewhere, didn't match the model's expectation for a business like yours. A few things make the difference between a quick resolution and a drawn-out one:
- Respond in writing, and respond fast. Given the longer wait times for human review described above, an early, complete written response is more likely to get resolved on the first pass than a phone call that gets you back in a queue.
- Lead with the reconciliation, not the explanation. Show the underlying transactions that produced the number in question before you explain why they're legitimate. Examiners move faster when the math is already done for them.
- Don't let one flagged year become three. An unresolved flag on this year's return is exactly the kind of anomaly that gets a follow-up model to take a second look at last year's and next year's filings too. Close it out completely rather than letting it linger.
- Bring in a CPA or enrolled agent if the notice mentions an expanded scope. A flag on one line item is manageable to answer yourself; a request for supporting records on multiple years is worth professional help.
What Actually Reduces Your Risk
None of this means you should under-claim legitimate deductions — the goal is a return that's accurate and well-supported, not a smaller one.
Keep contemporaneous records, not reconstructed ones. Documentation created at the time of the expense — a receipt logged the week you bought it — carries far more weight than a spreadsheet assembled the night before your return is due. It's also just easier: reconstructing six months of expenses from memory in April is how legitimate deductions get missed or misstated.
Use a chart of accounts that mirrors how your business actually runs. Broad, vague categories like "Miscellaneous" or "Other Expenses" are exactly the buckets that are hardest to defend and easiest for a model to flag as noise. If your books can't explain a number, neither can you.
Reconcile income against every 1099 and K-1 before you file. Since the IRS already has these forms, any mismatch is close to a guaranteed flag. A quick cross-check against your books before filing catches this in minutes instead of months later in a notice.
Avoid round numbers on anything you didn't literally spend in round numbers. If your real bookkeeping is precise, your return will naturally reflect that precision — which is itself a signal of legitimacy.
Separate business and personal accounts completely. Commingled funds are one of the fastest ways to turn a routine review into a real examination, because they make every other number on the return harder to verify.
The common thread: every one of these fixes comes down to keeping clean, complete, contemporaneous books throughout the year — not scrambling to reconstruct them at tax time. Plain-text, version-controlled accounting makes that discipline easy to sustain, because every transaction is recorded as it happens, in a format you can audit yourself before the IRS ever does.
Keep Your Books Ready for Any Kind of Scrutiny
As audit selection gets faster and more automated, the businesses with the least to worry about are the ones whose books were accurate before an algorithm ever looked at them. Beancount.io gives you plain-text accounting with full transparency and a complete, git-style history of every change — no black-box categorization, no reconstructed records, just a ledger you can defend line by line. Get started for free and keep your books audit-ready by default, not by scramble.