Most integration projects are scoped around the happy path. The order comes in from the website, lands in the ERP, creates the invoice, updates the CRM. Everyone tests that path, everyone signs off, and everyone moves on.
Then a customer enters a postal code the ERP validator rejects. A product gets archived while an order for it is mid-flight. An API times out during a vendor's maintenance window. The record doesn't process. It goes somewhere — a log file, an error table, a dead-letter queue — and it sits there. Nobody is assigned to look. Weeks later, someone in finance asks why three orders never invoiced, and the trail leads to a folder of failures nobody knew existed.
This is the most common integration failure mode we see, and it has almost nothing to do with the quality of the code. It's an ownership gap.
Failures Are Normal; Silence Isn't
Start by accepting that a percentage of records will always fail. Systems go down. Humans type things. Vendors change field requirements without telling you. An integration that never errors is usually an integration that's silently dropping data.
The goal isn't zero failures. The goal is that every failure is visible, categorized, and resolved inside a known window. That's a business process, not a technical feature — and it's the part that gets cut when a project runs long.
Separate the Three Kinds of Failure
Lumping all errors together is why error logs get ignored. Ninety percent of entries are noise, so the real problems get buried. Sort them:
- Transient — timeouts, rate limits, brief outages. These should retry automatically and never reach a human.
- Data — a missing tax ID, an unmatched SKU, a customer record that doesn't exist yet downstream. A human needs to fix the source data, then reprocess.
- Logic — the integration is doing the wrong thing. A mapping is stale, a business rule changed, a new product type isn't handled. This needs a developer, not a data fix.
Once you split these, the picture changes. Transient errors disappear into automated retries. Data errors go to the operations person who owns that record type. Logic errors become a small, honest backlog of work. Suddenly the queue is short enough that someone will actually look at it.
Retries Need Rules, and Rules Need Idempotency
Automatic retries are the right answer for transient failures, but only if reprocessing the same record twice can't create two invoices. Idempotency — the property that running an operation again produces the same result, not a duplicate — is something you design in, usually by having the receiving system check a unique external reference before creating anything.
If your integration isn't idempotent, aggressive retries make things worse. Fix that first, then set a sane policy: a handful of retries with increasing delays, then stop and escalate. Retrying forever just hides the problem in a different place.
Give the Queue a Name and a Face
Here is the part that actually solves the problem. Someone owns the exception queue by name. Not "IT." Not "the vendor." A person, with a backup, and a stated expectation for how quickly items get triaged.
In most mid-market companies that person shouldn't be a developer. Data errors are business decisions — which customer record is correct, whether to accept the order, what the right account is. An operations lead with access to a clear exception screen will resolve most of the queue faster than an engineer digging through logs.
That means the exception view has to be built for a human. Plain language description of what failed, which record, what field, and a button to reprocess after the fix. If your only interface is a raw log, you've guaranteed that only engineers can help.
Review the Patterns, Not Just the Items
Clearing the queue every day is maintenance. Reading the queue every week is improvement.
Spend twenty minutes looking at what failed and why. The same three error types usually account for the bulk of the volume, and most of them are fixable at the source: tighten a form validation on the website, add a required field in the CRM, build a lookup for the mismatched product codes. Each fix permanently removes work from someone's day.
This is also how you catch the slow disasters. A steady trickle of failures from one region, one product line, or one integration partner is a signal long before it becomes an escalation.
What Good Looks Like
You'll know the exception desk is working when someone can answer three questions without opening a ticket: how many records failed yesterday, who is fixing them, and what the oldest unresolved item is. If those answers require an investigation, you don't have an integration problem — you have an unowned one.
If you're connecting an ERP, a CRM, a website, and a handful of tools and nobody's quite sure where the failures land today, that's a conversation worth having. We build these systems — and the exception handling around them — for a living. Reach out to Infraxio and we'll walk your data flows with you.
