# SMS Queued-State Incident — 14 July 2026

import { internalAllowedTenants } from "../../src/utils/allowed-tenants";
import { Head } from "zudoku/components";
import { Callout } from "zudoku/ui/Callout";
import RequireTenant from "../../src/components/RequireTenant";

<RequireTenant allowedTenants={internalAllowedTenants}>

<Head>
  <title>SMS Queued-State Incident — 14 July 2026 — Tech231 Platform</title>
  <meta name="description" content="Database-backed RCA for the 14 July 2026 SMS dispatch outage." />
  <link rel="canonical" href="https://docs.tech231apps.net/internals/sms-queued-state-incident-2026-07-14" />
</Head>

## Summary

On 14 July 2026, outbound SMS messages were accepted into SMS state storage but did not progress beyond `Queued`. The stranded-message set was created from **00:05:23 UTC** through **20:57:46 UTC**. The high-volume retry surge began at **19:45 UTC**. At the time of investigation on 15 July, **1,931 distinct messages** from that set remained queued.

The database evidence shows repeated internal attempts before SMSC connector selection, rather than provider-side rejections. Normal submission of newly created messages resumed at **21:22:24 UTC** after the LCC SMPP connectors reconnected.

<Callout type="caution" title="Backlog remains stranded">
The 1,931 affected messages were not present in the durable connector queues and have no recorded SMSC assignment. Restoring connector health alone did not retry or deliver this backlog.
</Callout>

## Scope and evidence

This RCA is based on read-only production PostgreSQL queries against `smsDb` on 15 July 2026:

- `sms.mt_doc_sms_states` — authoritative current SMS lifecycle state.
- `sms.message_history_projection` — denormalized lifecycle projection.
- `sms.mt_doc_sms_connectors` — persisted SMSC connector state.
- `sms.mt_doc_queued_sms_messages` — durable per-connector queue payloads.

No recipient numbers, message content, or tenant identifiers were queried or recorded in this document.

## Customer impact

| Metric | Observed value |
| --- | ---: |
| Messages created on 14 July | 2,007 |
| Distinct messages still `Queued` | 1,931 |
| First affected message | 00:05:23 UTC |
| High-volume retry surge begins | 19:45 UTC |
| Last affected message | 20:57:46 UTC |
| Largest one-hour intake | 1,213 messages, 16:00–16:59 UTC |
| First post-recovery provider submission | 21:22:24 UTC |
| Successful delivery receipts for messages created after recovery | 62 |

The affected records have no provider submission timestamp, no selected SMSC configuration ID, and no failure reason. They did not reach provider submission.

## Timeline

All times are UTC.

| Time | Observation |
| --- | --- |
| 00:05:23 | First message that remains queued was created. |
| 00:00–15:59 | Every message created in this window remains queued; no provider submissions are recorded. |
| 16:00–16:59 | 1,213 messages were created and remained queued, creating the principal backlog. |
| 19:45 | High-volume retries began: 2,066 retry events across 423 messages in the next five minutes, compared with 194 events in the preceding five minutes. |
| 20:57:46 | Last message in the stranded group was created. |
| 20:58:58 | Last recorded state update for any stranded message. |
| 21:22:24 | First provider submission for newly created traffic appears in the history projection. |
| 21:22:24–22:15:57 | Ten LCC SMPP connectors recorded reconnection times. |
| 21:22–23:59 | New traffic resumed normal lifecycle activity: 62 delivered, 1 sent, and 13 failed. |

## Technical findings

### The failure occurred before connector selection

All 1,931 stranded records contain `LastAttemptAt`, but none contains:

- `SmscConfigIdUsed`
- `SubmittedAt`
- `FailureReason`

This rules out a normal provider submission failure for the affected messages. The dispatch path was attempting work but did not assign an SMSC connector or submit to a provider.

### Retry behavior was unsafe

The affected messages recorded the following retry-attempt distribution:

| Retry attempts | Messages |
| --- | ---: |
| 0 | 20 |
| 1–9 | 16 |
| 10–99 | 869 |
| 100–999 | 999 |
| 1,000+ | 27 |

The highest observed retry count was **1,226**. This indicates a tight or ineffective retry loop. The state remained `Queued` instead of transitioning to a diagnosable terminal failure, retry-with-backoff state, or assigned connector queue.

### Connector recovery did not drain the backlog

At investigation time, all ten LCC SMPP connectors persisted a `Connected` state. Their reconnect timestamps cluster from 21:22 through 22:15 UTC, matching the return of normal submissions for new traffic.

However, `sms.mt_doc_queued_sms_messages` contained only seven current durable queue items. Five were failed entries from the recovery window with `SMPP client is not connected.` The 1,931 stranded state records were not in a connector queue, so connector recovery could not drain them.

```mermaid
flowchart LR
  A[SMS accepted] --> B[SMS state persisted as Queued]
  B --> C{SMSC selected?}
  C -- No: affected records --> D[Repeated internal attempts]
  D --> B
  C -- Yes: normal path --> E[Connector durable queue]
  E --> F[Provider submission]

  style C fill:#fef3c7,stroke:#b45309
  style D fill:#fee2e2,stroke:#b91c1c
```

## Root-cause statement

**Proximate cause:** The SMS dispatch/routing path repeatedly attempted affected states but did not select an SMSC connector. This left messages in `Queued` without provider submission or a durable connector-queue handoff.

**Root cause:** Not yet established from database state alone. The database does not preserve the exception, configuration event, grain activation, or deployment action that caused connector selection to fail. Logs and traces from **00:05–21:22 UTC** are required to identify that initiating fault.

**Contributing design gap:** The retry path allowed very high retry counts while preserving `Queued` status and omitting a structured failure reason. This obscured detection and left the backlog without an automatic recovery route once connectors were healthy.

## Recovery

The observed recovery point is the first normal submission at 21:22:24 UTC and the associated LCC connector reconnects. The production data does not establish which operator action, deployment, configuration change, or automated process initiated that recovery.

The existing backlog was not recovered automatically. It requires an explicit, idempotent replay or repair procedure that first prevents duplicate provider submission.

## Corrective actions

1. **Preserve evidence:** retrieve and retain SMS API, Orleans silo, and SMSC connector logs/traces for 00:05–21:22 UTC on 14 July.
2. **Repair the backlog safely:** build an idempotent operator procedure that identifies queued messages without `SmscConfigIdUsed` or `SubmittedAt`, re-runs connector selection, and records the selected connector before submission.
3. **Bound retries:** enforce exponential backoff and a terminal/dead-letter transition with a structured reason when connector selection repeatedly fails.
4. **Alert on dispatch invariants:** alert when queued states have no connector assignment, when retry counts exceed a bounded threshold, and when queued age exceeds the delivery objective.
5. **Make recovery measurable:** expose counts for state-queue backlog, connector-queue backlog, connector-selection failures, retry bands, and oldest queued age.
6. **Test the failure path:** add an integration test where no connector is available or connector selection fails, verifying bounded retries, an observable failure reason, and subsequent recovery after connector restoration.

## Follow-up verification

Before closing the incident, verify that:

- The 1,931-message set has been handled by an approved replay, expiry, or customer communication process.
- No messages remain queued without `SmscConfigIdUsed` beyond the configured threshold.
- Retry counts cannot increase indefinitely.
- Alerts fire for a simulated connector-selection failure.
- The root initiating event is confirmed from logs or traces and this RCA is updated.

## Related pages

- [SMSC Connection Manager Service](/internals/smsc-connection-manager-usage)
- [Sender ID Operations Guide](/senderid/operations-guide)

</RequireTenant>
