Canopus Software & Engineering — home

Why integrations fail silently

A broken screen is reported in minutes. A sync that stopped writing to the ERP three weeks ago is discovered at stocktake — after twelve hundred orders reconciled wrong. That gap is what makes integration failures expensive.

Why integrations fail silently — Canopus engineering guide
The short answer

Integrations fail silently because nothing is watching for absence. Monitoring catches errors, but a stopped sync produces no errors — it produces nothing at all. The fix is to alert on missing expected traffic and run a scheduled reconciliation that proves both systems still hold the same records.

Five ways integrations go wrong

Every integration we have been called in to rescue failed in one of these five ways. Each has a known engineering answer — the only question is whether it was built in at the start or left out.

1. The double-processed webhook

The sender times out waiting for your response and retries. Your system creates the order — or issues the refund — twice. This is not an edge case: Stripe, Shopify and most carrier APIs retry by design, and a slow database query is enough to trigger it.

The fix: idempotency keys, recorded before processing. Every inbound event carries a unique identifier that you store the moment it arrives. A repeat delivery is recognised and acknowledged without re-running the work. This is a handful of lines and it eliminates an entire class of duplicate-record bugs.

2. The silent stop

A token expired in March. Nothing alerted, because nothing failed loudly — the integration simply stopped being called. Nobody noticed until someone asked why April's numbers looked light.

The fix: heartbeat monitoring plus scheduled reconciliation. Alert on the absence of expected traffic, not just on errors. Then run a nightly job that compares record counts and checksums on both sides. Silence is not evidence that data is flowing; a reconciliation job that finds nothing is.

3. The partial write

An order is created in system A, the call to system B fails, and now the two disagree with no record of why. The user saw a success message. Support finds out a fortnight later.

The fix: a durable queue with dead-letter handling. The event survives the failure and retries with exponential backoff. Anything that still cannot be processed lands in a dead-letter queue that a human can actually see and act on — not a log line nobody reads.

4. The version that moved

The vendor deprecated v2 of their API. The notice went to an inbox belonging to someone who left in 2024. The integration worked right up until the sunset date, then stopped.

The fix: pin the API version explicitly, run contract tests nightly against the live sandbox, and route vendor notices to a team address. Contract tests catch a changed response shape days before it reaches production.

5. The field that meant two things

"Reference" is the purchase order number in one system and the invoice number in the other. Both teams are certain they are right. The data flows perfectly and is wrong in a way no error can detect.

The fix: a written field map agreed before any code. Which system is authoritative for each field, and what happens when they conflict. Most integration failures are ownership disputes, not technical faults.

The common thread

Four of these five produce no error at all. If your integration monitoring only watches for exceptions, it is watching the one failure mode that rarely happens.

Choosing the right pattern

Picking the wrong integration pattern is how a simple connection becomes a permanent maintenance cost. If you are still budgeting a product rather than repairing a live one, what it has to talk to is one of the four decisions that drive the cost of a SaaS MVP, so the choice below is a commercial one as much as a technical one. Four patterns cover almost everything.

PatternUse whenAvoid when
Synchronous API callThe user is waiting and needs the answer now — card authorisation, stock check at checkoutThe other system is slow or unreliable; you are now only as available as they are
Webhook + queueThe other side pushes events — payments, shipment status, form submissionsStrict ordering matters and the sender does not guarantee it
Scheduled batchHigh volume, no urgency — nightly catalogue, daily financial posting, EDI runsUsers expect near-real-time and will call support about it at 11am
Event bus / streamingSeveral systems need the same event and the list will growYou have two systems and no plans for a third — it is overhead you maintain forever

Scroll the table sideways for the full comparison.

The mistake we are most often called in to undo

A synchronous call to a third-party API inside the checkout path. It works perfectly in testing. Then the provider has a slow morning and your checkout starts timing out alongside it.

Anything not strictly needed to answer the user belongs in a queue. The order is accepted, the downstream write happens behind it, and a failure there becomes an alert rather than a lost sale. This single change has more impact on revenue than any amount of retry tuning. The queue is not free infrastructure, though: a managed broker, the workers draining it and the log volume they produce become a permanent line on the monthly bill, which is worth sizing deliberately rather than meeting at quarter end — where cloud spend actually leaks covers what that line tends to look like.

What a healthy integration ships with

  • A dashboard showing throughput and failure rate, visible to operations rather than only to engineers
  • Alerting on failure rate and on queue depth and on unexpected silence
  • A dead-letter queue with an owner and a review routine
  • A scheduled reconciliation report proving both sides agree
  • A written runbook for the top three failure modes
  • Pinned API versions with contract tests running nightly

If an existing integration is missing these, the first step is not to fix the code — it is to add the visibility, so you can see the real failure rate before and after any change. That approach is described on our API and system integration page.

Questions

What teams ask before wiring two systems together

How long does an integration take, and what does the other side's sandbox do to that estimate?

For a documented REST API with a working sandbox, a single flow is usually two to three weeks of engineering including reconciliation and alerting; four to six where both systems write and conflicts have to be resolved. Few engagements are one flow, which is why the range we publish is four to twelve weeks. The estimate that moves is rarely ours. Sandbox credentials from a bank, an ERP vendor or a carrier can take weeks to arrive, arrive without the endpoints you actually need, or hold data that behaves nothing like production. We request access on day one, sequence the build so nothing idles while we wait, and mark in the estimate which dates depend on somebody else's queue rather than on us.

The vendor's API has no sandbox at all. What then?

You build against a recorded contract and go live with the smallest blast radius you can arrange. Capture real responses once by hand, freeze them as fixtures, and replay them in CI through a mock server such as WireMock or Prism, so a changed response shape still fails a build. Then ship the write path behind a flag in dry-run mode — it logs the exact payload it would have sent, without sending it — and enable it for one ring-fenced account, low values, reversible operations first. We will not point a first-time write at a live endpoint with no rollback path and nobody watching.

Who owns the field map, and how do conflicts get resolved?

You own it. Which system is authoritative for a customer address or an order reference is a business decision, and no engineer can settle it by preference. We produce the map in discovery: every field, its owning system, its direction of travel, and what happens on conflict. Three rules cover most of it — source of truth wins, last write wins, or quarantine for a human. Money, tax and identity fields never get last write wins; a mismatch parks the record in an exceptions queue with both values visible. One named person signs the map before we build, because the alternative is finding the disagreement in production.

What does a reconciliation job actually prove, and how often should it run?

It proves both systems hold the same records, with matching values in the fields you agreed to compare, across a defined window — nothing more. It cannot tell you the mapping is right: if "reference" means two different things, both sides will agree on a value that is wrong. Nightly over a rolling 48 hours is the default, so a late write is not reported as a discrepancy. Hourly where a day of drift is expensive: payments, or stock in a warehouse promising same-day dispatch. Alert on a mismatch, and separately on the job failing to run — a reconciliation that quietly stopped looks exactly like a clean one.

What happens when the vendor deprecates the API version we are on?

Vendors typically give six to twelve months, though the notice period is theirs to set and some give considerably less. It stays scoped work rather than an emergency provided somebody saw the notice. Vendor changelogs go to a team address that survives a resignation, the version stays pinned in configuration rather than tracking whatever is current, and a nightly contract test runs against the sandbox on the next version — so a changed response shape fails a build well before the sunset date. The cost splits cleanly: a renamed field is hours, a reshaped webhook payload or a new authentication model is weeks. The small end comes out of the monthly enhancement hours on an AMC support retainer; a reshaped payload is more than a month's allowance covers, so it is quoted before it starts, the same as on a fixed-price build. We would rather say that now than let you assume a migration is included.

Is an integration platform (iPaaS) worth it, or should we write the code?

Count the flows and look at their shape. A platform — Workato, Boomi, MuleSoft, Azure Logic Apps — earns its licence when you have many small flows between mainstream SaaS products that already have connectors, and the people who own the mapping are not engineers. Writing the code wins when the logic is specific to your business, when idempotency and reconciliation have to behave exactly as you decided, or when per-task pricing at your volume overtakes the cost of maintaining it. The trap sits in the middle: the connector that does most of the job and cannot be extended for the rest. What we build is the code path. We resell nothing and earn no commission from any vendor, so where the flow count and shape point at a platform we will say so and give you the criteria to pick one, rather than quote you a build.

Published: Last updated: Written by Canopus delivery team

Get a Free Quote

Which two systems disagree, and who re-keys the difference?

That is the whole brief. Tell us the systems and the field they argue about, and an engineer replies within one business day with the pattern we would use.

Scope an integration

One business day, from an engineer who has integrated these before.
Call WhatsApp Get a Quote