01Built-in controls for dependable data flows
| Control | Operational benefit |
|---|---|
| Policy-driven recovery | Flower combines retry classification, backoff and circuit policies to automate recovery within the data flow. Execution context and proactive alerts give operators a clear view of progress and guide the next action. |
| Completion verification | Combine completion checks for file-copy routes with schema, record-count and business-total validation for processed data. Acceptance criteria become part of the flow, with delivery evidence available to operations. |
| Governed database operations | Bring queries, prepared writes and explicit transactions into declarative data flows. Database permissions govern access, while validation, reconciliation and execution records connect the movement of rows to your operating policies. |
Failure classification
02Decide what failed before deciding what to repeat.
Flower classifies transport conditions, data-quality results and destination outcomes to select the appropriate operating response. Recovery, reconciliation, quarantine and alerts work together within the same flow.
Transient and retryable
Flower responds to temporary unavailability with controlled retries, backoff and jitter. Recovery progress, backlog and alerts remain visible in the same operating view, helping teams keep recurring deliveries on track.
Permanent and actionable
When input or access needs attention, Flower isolates the affected unit and preserves the execution context. Configured quarantine and notification policies direct the information to the responsible team for a focused response.
Uncertain and reconcile-first
A connection can fail after the destination committed. Flower checks destination identity, tracking evidence, counts, or checksums before replay so an unknown outcome does not automatically become duplicated work.
Recovery control loop
03Advance only after the unit is safe and evidenced.
The same control loop follows data from source to destination. Routine transient recovery runs automatically; unsafe or uncertain outcomes retain context, follow configured policy, and raise an actionable signal only when operator judgment is needed.
Choose the restart boundary
Use an object, file, deterministic row batch, or other smallest safe unit. Store the checkpoint outside transient execution state and advance it only after destination confirmation.
Gate progress with validation
Transport completion is one signal. Structure, schema, counts, business rules, integrity, and destination reconciliation decide whether the unit may be published, retained for review, or rejected.
Keep recovery evidence attached
Record attempts, classification, source and destination identity, validation outcome, quarantine, reconciliation, and final action in one lineage path. Alerts carry enough context for an operator to act.
Shaped by production
04Reliability earned under real data pressure.
Flower has evolved around continuous production workloads since 2019, including very large telecommunications volumes and business-critical financial data. That operating experience informs its integrated approach to recovery, validation and visibility.
Failure is explicit
Incomplete transfers, invalid records, and policy violations become visible states with defined handling instead of silent downstream defects.
Recovery follows your policies
Retry budgets, backoff, restart boundaries, quarantine, destination reconciliation, and lifecycle actions adapt to the endpoint and workload rather than being improvised during an incident.
Evidence stays attached
Lineage, metadata, events, metrics, and alerts help operators reconstruct what moved, what changed, and what needs attention.
Recovery acceptance
05Specify recoverability before the first incident.
Evaluate self-healing on your own workload with clear delivery criteria, recovery policies and operational evidence. Repeat these checks as endpoints and policies evolve to carry a dependable operating model from evaluation into production.
Write the recovery contract
For each unit, state permitted loss, duplication, ordering change, maximum delay, and required destination evidence. A precise contract prevents teams from calling a restarted process successful while the data outcome remains unknown.
Inject distinct failure classes
Test source unavailability, transport interruption, throttling, invalid credentials, schema rejection, corrupt payload, full local storage, and target timeout separately. Each should reach the intended retry, stop, quarantine, or reconciliation path.
Force an ambiguous commit
Drop the connection after the destination may have committed but before confirmation reaches the sender. Verify that destination identity, counts, checksums, or another durable signal are checked before any replay is allowed.
Restart across the state boundary
Restart the worker, host, and dependent state service at different points. Confirm that checkpoints, pending attempts, quarantine context, and lineage survive together, and that an upgrade or rollback does not reinterpret in-flight work.
Measure pressure and catch-up
Hold an endpoint unavailable long enough to build a realistic backlog, then restore it. Measure queue age, disk growth, retry rate, useful catch-up throughput, resource saturation, and whether new traffic is starved or remains protected.
Prove quarantine and replay
Send invalid and policy-violating units beside valid work. Confirm that unsafe data stops without blocking unrelated progress, keeps enough context for diagnosis, and can be corrected and replayed without bypassing validation or duplicating completed units.
Prove durable-state restoration
Remove or corrupt the active state store and restore it from the documented backup or replica. Measure the recovery-point gap, reconcile source and destination evidence, and confirm that the rebuilt checkpoint neither skips eligible work nor silently republishes completed units.
Test the operator handoff
Exhaust a retry budget and trigger a permanent failure. The alert should identify the affected unit, source and destination, last safe boundary, attempts, evidence, owner, and approved next actions without requiring reconstruction across unrelated tools.
Gate releases with recovery drills
Keep the failure matrix, expected evidence, and recovery timings as a repeatable suite. Run it after runtime, connector, schema, policy, or infrastructure changes so recoverability is a tested release property rather than historical confidence.