Lighthouse · Cases

From a support ticket to a bad-call attribution

How the Cases module turns one customer complaint into call telemetry — and how doing that on ticket 1376296 surfaced (and then sized) a kind of “no-ring” that isn’t a quality defect.

Botim VoIP Qualitymodule: /casesexample: ticket 1376296Prepared by Zhifeng Guo

What Cases is

Cases is the bottom-up half of Lighthouse: start from a real signal — a CS ticket or a low-star rating — and walk down to the telemetry that explains it. It resolves the complaint to a user, pulls that user’s recent VoIP calls straight from the footprint, and lets you read each call the way the client experienced it.

The hard part it hides is identity: a CS ticket keys on a support user-id, the footprint keys on the MSISDN, ratings key on yet another id. Cases bridges them (phone-in-text → author history → room-id) so “this complaint” and “these calls” line up.

▸

The workflow

Six steps, top of the funnel to root cause.
1

Open the ticket

Go to /cases/<ticketId>. Cases fetches the CS ticket, extracts the phone, and resolves it to the footprint user. You see the complaint text, category, and a first machine verdict.

2

Read the profile bar

One glance frames the ticket: call volume, accept rate, and what the calls cluster on — device, client version, relay, ISP — plus the top end-reasons and red-flag callouts (mostly-unconnected, single-version, send-mute, retry storm).

3

Scan the call list

One row per call: outcome, end-reason as rN · label, RTT / loss, relay, device, and the user’s own ★ rating. Bad values are flagged so the problem calls stand out.

4

Expand a call

Full per-call detail — audio internals (PLC / jitter-buffer / quality head→tail), video, transport depth, and permissions (mic / camera). Each reason shows its label and the fingerprint evidence. “Compare both sides” puts the user’s leg next to the peer’s to separate server/relay problems from local network ones.

5

Trace the server signaling

For “can’t connect / no-ring” calls the client can’t tell you why — so follow the call’s traceId into the server call-trace: S1 chat-server start → S2 im-push → S3 FCM/APNs. The stage that fails names the culprit.

6

Hand off to Sherlock

Copy or download the whole context — telemetry, reason labels, identifiers — and continue the investigation in the Sherlock debug agent.


▸

Worked example — ticket 1376296

“Botim not working. The call is getting disconnected the minute I am calling.”
#1376296Botim not working Voice call issueIndia · Airtel · WiFi iPhone 15 Pro · iOS 18.3.2 · v4.14.0verdict: no-ring, not quality

Reading the profile bar and drilling one failed call produced a single, clean chain of evidence:

  • Complaint
    User can’t make calls — every call drops “the minute I am calling.”
  • Profile · 99 calls
    3% accepted · 93 lasted ≤5s · 89× end-reason r44
    Not a flaky-quality pattern — the calls die at setup, before media.
  • Reason decode · r44
    outbound setup failed — no room
    Fingerprint over 8 days / 429M rows: r44 is 100% caller · 0% ring · 100% empty roomId — the only value that never gets a room.
  • Server trace · S1
    Get_Room_Failed (S2 / S3 never reached)
    The chat-server refused to allocate a room for this user’s outbound call — so it never rings. Independent confirmation of the r44 fingerprint.
  • Root cause
    The user is blocked by our risk-control policy.
    Room allocation is being denied by design — this account is under a risk-control / blocking rule. The failure is a policy action, not a network or client fault.
SignalValueReads as
Calls in window99whole complaint window
Accepted3 · 3%almost nothing connects
End-reason r4489 · 90%outbound setup failed
Empty roomId on r44100%never got a room
Server S1Get_Room_Failedserver refused the room
↗ Open this case in Lighthouse — /cases/1376296 internal network · access code required
The expanded call detail for #1376296 — reason r44, family cards all OK (no media), the “send-mute@pre (recVol≈0)” signal, and the identifiers (roomId / traceId / uid / peerUid). Click to open full size.
The finding · bad-call attribution

Risk-control blocks are being counted as “no-ring” bad calls.

These calls never ring because we denied the room — a deliberate risk-control action — yet they sit in the same no-ring bucket that feeds the bad-call rate. The metric is crediting a policy action as a quality failure. Blocked ≠ broken.

Sized fleet-wide over Aug 4–10: excluding risk-control-blocked no-ring moves the total bad-call rate by up to 0.6% depending on definition — a real, systematic shift in a rate this closely watched.

bad-call rate drop if blocked no-ring is excluded  (Aug 4–10)
vs A/V/network bad    0.0%
vs no-ring alone      0.6%
vs van + no-ring      0.3%

Cross-check: r44 is ~1.7% of never-ring calls, and risk-control blocks are ~0.35–0.42% of all no-ring — a systematic miscount, not noise. It moves the number and it is the wrong number to move.

▸

What to do

01

Re-bucket the metric

Report risk-control Get_Room_Failed separately from no-ring, so the bad-call rate counts genuine quality failures — not policy actions we took on purpose.

02

Keep the fleet split live

The Aug 4–10 sizing already exists as a widget; keep r44 / Get_Room_Failed split from genuine setup failures so the share is watched, not re-derived.

03

Close the CS loop

These users are blocked by design. CS should see a “restricted by policy” state instead of chasing a call-quality bug that isn’t there.

04

Reuse the method

The value here is the workflow: ticket → r44 → Get_Room_Failed → policy, then size it before acting. Same path fits the next attribution question.

Sources & caveats. Call telemetry: ES footprint logs-footprint-voipstat-* (≈7-day retention). Server signaling: the server call-trace (StarRocks), joined by traceId. Reason-code labels (rN) are behavioral-fingerprint inferences over an 8-day / 429M-row window — there is no official BRtcKit enum; treat labels as strong hypotheses, not spec. “Risk-control block” is the root-cause reading of a confirmed Get_Room_Failed; confirm against the blocking-policy system for the exact rule.