SAP Integration Suite Failover Tenant Sync: How to Keep Primary and Failover Tenants in Parity Without Manual Redeployment

A Failover tenant protects the business only if it matches production at the moment of cutover. You keep the two tenants in parity by syncing artifacts continuously, checking deployment status and deployed versions per artifact on both sides, recording a deployment mode (Single or Both) for each artifact, and running failover in a fixed priority order. Manual redeployment before each drill cannot keep up with a landscape that changes every week. Drift then stays invisible until the day you need the Failover tenant.
This article explains how to prove Failover readiness at the artifact level. For the wider question of what SAP provides natively for high availability and what you must build for cross-region disaster recovery, see our analysis of SAP Integration Suite high availability vs disaster recovery.
How does a Failover tenant drift out of sync?
A Failover tenant drifts when a change reaches the Primary tenant through a path that does not also update the Failover tenant. In most multi-tenant landscapes, five causes explain nearly all drift:
- Hotfixes deployed directly on Primary. A developer fixes a production mapping under time pressure, and the change never goes through the transport route that feeds the Failover tenant.
- Configuration-only changes. An externalized parameter, such as a receiver URL or a timeout, changes on Primary. The iFlow version number stays the same, so a version check alone does not catch it.
- Security material rotation. A certificate or password is renewed on Primary only. The iFlow on Failover deploys without errors and then fails on its first call.
- Partial redeployment. A package is redeployed, but one artifact in it fails to start and nobody checks it.
- Runtime state left behind. A scheduled flow stays active on Failover after a drill, or a number range on Failover still holds last month's value.
None of these show up in tenant-level health checks. Both tenants are reachable, and both report as healthy. The gap only appears when traffic arrives.
What must a Failover readiness view show?
A readiness view must answer one question: if we fail over now, will the correct artifacts run, in the correct version, in the correct place? To answer it, the view needs landscape context plus four artifact-level checks.
The context comes first. The view shows:
- Which tenant is Primary and which is Failover
- Which runtime belongs to each tenant
- Whether the landscape is in Normal or Failover state
This matters in hybrid landscapes that mix the SAP Cloud Integration runtime with Edge Integration Cell runtimes on customer-managed Kubernetes. In those landscapes, one tenant can map to several runtimes on the Failover side.
┌──────────── Routing layer (health-checked) ────────────┐
│ │
Normal state ▼ Failover state ▼
┌──────────────────────────┐ continuous sync ┌──────────────────────────┐
│ PRIMARY TENANT │ ────────────────────▶ │ FAILOVER TENANT │
│ Runtime: Cloud / Edge │ design-time + │ Runtime: Cloud / Edge │
│ Group health probe ● │ runtime artifacts │ Group health probe ○ │
│ Single-mode flows: ON │ │ Single-mode flows: OFF │
│ Both-mode flows: ON │ ◀── parity checks ──▶ │ Both-mode flows: ON │
└──────────────────────────┘ status + version └──────────────────────────┘
How do failover groups and subgroups control sequencing?
Failover groups and prioritized subgroups make the cutover order deterministic. Each subgroup has a priority, and the tooling moves artifacts in that order. The team decides the sequence before an incident, not during one.
Each failover group links these elements:
- A Primary tenant and runtime
- A Failover tenant and runtime
- A group health probe that shows which side is active
- A set of artifacts split into prioritized subgroups
A typical priority order looks like this:
| Priority | Subgroup | Why it goes in this position |
|---|---|---|
| 1 | Shared utilities, value mappings, script collections, credentials | Other flows call or read them at runtime |
| 2 | Process Direct and JMS consumer flows | Receivers must be ready before senders start to deliver |
| 3 | Inbound listener flows (HTTPS, SOAP, OData, IDoc) | These accept traffic after the routing switch |
| 4 | Scheduled and polling flows | These start last, so they pull data only when every downstream flow is ready |
Because the groups follow the real dependencies between interfaces, the order you rehearse in a drill is the same order that runs in a real event. The failover plan lives in the tooling and not in a document.
Why check SAP iFlow deployment status per artifact?
Tenant health tells you that a tenant is reachable. Artifact status tells you that a specific iFlow can process messages. You need the second view on both sides to find sync gaps before cutover.
SAP Cloud Integration reports four runtime states for a deployed artifact. A comparison tool adds two more states when it finds an artifact on one tenant only:
| State | Source | What it means before cutover |
|---|---|---|
| Started | SAP runtime | Ready to process messages |
| Starting | SAP runtime | Deployment or failover is in progress |
| Stopping | SAP runtime | Undeployment is in progress, for example a Single-mode flow during failover |
| Error | SAP runtime | Deployment failed. On Failover, this is a defect that stays hidden until traffic arrives |
| Stopped | Derived | Deployed in design time but not running |
| Not Present | Derived | The artifact does not exist on this tenant. This is a sync gap |
Two combinations need action before any cutover:
- Started on Primary and Not Present on Failover. Sync the artifact before you continue.
- Error on Failover. Read the deployment error. The cause is usually a missing credential, key pair or referenced artifact.
The transitional states, Starting and Stopping, let you follow a failover or failback while it runs, instead of waiting to see whether it worked.
How do you validate iFlow version parity between tenants?
You validate version parity by comparing the deployed runtime version of each artifact on both tenants, and not only the design-time package version. An iFlow can be saved as version 1.0.7 in the package while the runtime still runs 1.0.6, because someone saved the change but did not deploy it.
The SAP Cloud Integration OData API exposes deployed artifacts with their ID, version and status. The following sketch shows the core of a parity check:
import requests
def runtime_state(base_url, token):
r = requests.get(
f"{base_url}/api/v1/IntegrationRuntimeArtifacts",
headers={"Authorization": f"Bearer {token}", "Accept": "application/json"},
timeout=30,
)
r.raise_for_status()
return {a["Id"]: (a["Version"], a["Status"]) for a in r.json()["d"]["results"]}
primary = runtime_state(PRIMARY_URL, primary_token)
failover = runtime_state(FAILOVER_URL, failover_token)
for artifact_id, (p_ver, p_status) in sorted(primary.items()):
f_ver, f_status = failover.get(artifact_id, (None, "NOT_PRESENT"))
if f_status == "NOT_PRESENT":
print(f"GAP {artifact_id}: missing on Failover")
elif p_ver != f_ver:
print(f"DRIFT {artifact_id}: Primary {p_ver} / Failover {f_ver}")
elif f_status == "ERROR":
print(f"DEFECT {artifact_id}: Error on Failover")
A version check catches changes to logic. It does not catch configuration-only drift. A complete parity check also compares externalized parameter values, and it applies the tenant-specific values that must differ between the two sides.
What automated sync covers
Continuous synchronization keeps these artifacts aligned between tenants:
- Design time: integration packages, integration flows, message mappings, value mappings, script collections, imported archives, schemas, message types, data types, service interfaces, custom integration adapters and integration flow configurations.
- Security and runtime: user credentials, number range objects, JDBC data sources, access policies and global variables.
What sync alone does not solve
Some items need design decisions in addition to sync:
- Tenant-specific configuration. Receiver hosts, Cloud Connector location IDs and OAuth token URLs can differ between regions. Mark these parameters as tenant-specific so that sync does not overwrite them.
- Secrets. Some secret values cannot be read back through the APIs. The sync process must store them securely, or you must maintain them on both tenants.
- Inbound trust. Sender systems must have valid credentials or client certificates for the Failover tenant before cutover. If they do not, the routing switch works, but every call gets a 401 response.
- Messages in flight. JMS queue content, Data Store entries and message processing logs stay on the Primary tenant. Define how you drain or replay them. The high availability vs disaster recovery analysis covers the options for runtime dependencies.
How do Single and Both deployment modes prevent duplicate processing?
The deployment mode of each artifact decides where it runs. Single means it runs on one tenant at a time. Both means it runs on both tenants all the time. When you record the mode per artifact, nobody has to remember which flows need special handling during failover.
| Mode | Typical artifacts | During failover | Risk if you set the wrong mode |
|---|---|---|---|
| Single | Timer-start flows, SFTP, mail and OData polling, batch extracts | Undeployed on Primary, then deployed on Failover | Two tenants poll the same folder or schedule, and files or records are processed twice |
| Both | HTTPS, SOAP, IDoc and OData listeners, called only through the routing layer | Stays deployed. The routing layer decides which tenant gets traffic | Low. A listener only processes the requests that are routed to it |
During a failover, Single-mode artifacts are stopped on Primary and started on Failover. The group health probe moves from the Primary runtime to the Failover runtime. Runtime values, such as global variables and number ranges, are synchronized as part of the same step.
Why number ranges must be current at cutover
Number ranges generate IDoc control numbers and EDI interchange control numbers. Suppose the Failover tenant resumes from a value synced a day earlier. It then issues numbers that trading partners have already received. Many partners reject duplicate interchange numbers, and some process the duplicates. For this reason, the number range sync must run at the moment of cutover, and not only on the regular sync schedule.
Post-mortem: what a DR drill found in a drifted landscape
The following composite scenario shows the drift pattern that parity checks catch. It is based on landscapes of about 400 artifacts across two regions.
A manufacturer ran a planned DR drill. It used a manual checklist and redeployed the key packages the week before. The drill found these problems:
- 7 artifacts were Not Present on Failover. All of them were added to Primary after the last full redeployment.
- 12 artifacts ran older versions on Failover. Two of them were mapping hotfixes that went to production directly.
- 3 artifacts were in Error on Failover because of a missing key pair alias after a certificate renewal.
- 1 SFTP polling flow was deployed on both tenants. Files that arrived during the drill were processed twice.
- 1 number range was behind. The Failover tenant issued EDI control numbers that a trading partner had already received.
The checklist showed that all packages were deployed. It did not show versions, per-artifact status or deployment mode. After the team moved to continuous sync with a daily parity report, drift was found within a day of each change and not at the next drill.
Failover readiness scorecard
Run this check before every maintenance window and DR drill. Any "No" blocks the cutover.
| Check | Pass criterion |
|---|---|
| Artifact presence | 0 artifacts Not Present on Failover |
| Deployed version parity | 0 version mismatches on artifacts in scope |
| Failover runtime health | 0 artifacts in Error |
| Deployment modes | All polling and scheduled flows set to Single |
| Number ranges and variables | Synced within the agreed cutover window |
| Inbound credentials | Sender access tested against the Failover endpoint |
| Subgroup order | Priorities reviewed after the last interface change |
Problem, solution and vision
Problem. Failover tenants drift through hotfixes, configuration changes and certificate renewals. Manual redeployment before each drill checks packages, not the runtime reality of each artifact.
Solution. Use continuous sync, per-artifact status and deployed-version checks on both tenants, a recorded Single or Both mode for each artifact, and failover subgroups with fixed priorities.
Vision. Parity becomes a gate in the delivery pipeline. A change to Primary is not complete until the Failover tenant shows the same version and state. Readiness then becomes a continuous, measured value that architects can report to the business at any time, and not a claim made once a year after a drill.
How Tarento closes the tenant-parity gap
Tarento's OneFailover automates the practices described in this article for SAP Integration Suite. It provides continuous artifact synchronization, per-artifact status and version parity across Primary and Failover tenants, deployment-mode enforcement, prioritized failover groups, and health-checked routing between regions.
iVolve, Tarento's SAP-certified integration framework, addresses the same problems from the delivery side. Its Transport and Governance Tool keeps versioning clean and standard, which stops unreviewed hotfixes from creating drift. Its Integration Inspector provides runtime monitoring and alerting, so an artifact that fails on either tenant is flagged before a cutover depends on it.
Frequently asked questions
How do I keep two SAP Integration Suite tenants in sync? Use continuous synchronization of design-time and runtime artifacts, not redeployment before each event. Then compare deployed versions and runtime status per artifact on both tenants to confirm that the sync worked.
What is version parity in SAP Integration Suite? Version parity means that each artifact runs the same deployed version on the Primary and Failover tenants. Compare the runtime version, because the version saved in the package can differ from the version that is deployed.
Which iFlows should not run on both tenants? Scheduled flows and polling flows, such as SFTP, mail and OData polling, should run on one tenant only. If they run on both tenants, they can process the same file or record twice.
Does artifact sync move messages in JMS queues? No. JMS queue content, Data Store entries and message processing logs stay on the Primary tenant. They need a separate drain or replay plan.
Why does an iFlow show Error on the Failover tenant after sync? The usual cause is a dependency that did not sync or that has a different value on that tenant. Examples are a credential, a key pair alias, a referenced script collection or a tenant-specific parameter.
How often should Failover readiness be checked? Check it continuously or at least daily, and always before a maintenance window or DR drill. Drift starts with the first change after the last check.
Is your Failover tenant ready for a cutover today? Talk to Tarento's Enterprise Integration team about a failover readiness assessment and a OneFailover setup for your SAP Integration Suite landscape.


