How the Cases module turns one customer complaint into call telemetry — and how doing that on ticket 1376296 surfaced (and then sized) a kind of “no-ring” that isn’t a quality defect.
Cases is the bottom-up half of Lighthouse: start from a real signal — a CS ticket or a low-star rating — and walk down to the telemetry that explains it. It resolves the complaint to a user, pulls that user’s recent VoIP calls straight from the footprint, and lets you read each call the way the client experienced it.
The hard part it hides is identity: a CS ticket keys on a support user-id, the footprint keys on the MSISDN, ratings key on yet another id. Cases bridges them (phone-in-text → author history → room-id) so “this complaint” and “these calls” line up.
Go to /cases/<ticketId>. Cases fetches the CS ticket, extracts the phone, and resolves it to the footprint user. You see the complaint text, category, and a first machine verdict.
One glance frames the ticket: call volume, accept rate, and what the calls cluster on — device, client version, relay, ISP — plus the top end-reasons and red-flag callouts (mostly-unconnected, single-version, send-mute, retry storm).
One row per call: outcome, end-reason as rN · label, RTT / loss, relay, device, and the user’s own ★ rating. Bad values are flagged so the problem calls stand out.
Full per-call detail — audio internals (PLC / jitter-buffer / quality head→tail), video, transport depth, and permissions (mic / camera). Each reason shows its label and the fingerprint evidence. “Compare both sides” puts the user’s leg next to the peer’s to separate server/relay problems from local network ones.
For “can’t connect / no-ring” calls the client can’t tell you why — so follow the call’s traceId into the server call-trace: S1 chat-server start → S2 im-push → S3 FCM/APNs. The stage that fails names the culprit.
Copy or download the whole context — telemetry, reason labels, identifiers — and continue the investigation in the Sherlock debug agent.
Reading the profile bar and drilling one failed call produced a single, clean chain of evidence:
| Signal | Value | Reads as |
|---|---|---|
| Calls in window | 99 | whole complaint window |
| Accepted | 3 · 3% | almost nothing connects |
| End-reason r44 | 89 · 90% | outbound setup failed |
| Empty roomId on r44 | 100% | never got a room |
| Server S1 | Get_Room_Failed | server refused the room |
These calls never ring because we denied the room — a deliberate risk-control action — yet they sit in the same no-ring bucket that feeds the bad-call rate. The metric is crediting a policy action as a quality failure. Blocked ≠ broken.
Sized fleet-wide over Aug 4–10: excluding risk-control-blocked no-ring moves the total bad-call rate by up to 0.6% depending on definition — a real, systematic shift in a rate this closely watched.
Cross-check: r44 is ~1.7% of never-ring calls, and risk-control blocks are ~0.35–0.42% of all no-ring — a systematic miscount, not noise. It moves the number and it is the wrong number to move.
Report risk-control Get_Room_Failed separately from no-ring, so the bad-call rate counts genuine quality failures — not policy actions we took on purpose.
The Aug 4–10 sizing already exists as a widget; keep r44 / Get_Room_Failed split from genuine setup failures so the share is watched, not re-derived.
These users are blocked by design. CS should see a “restricted by policy” state instead of chasing a call-quality bug that isn’t there.
The value here is the workflow: ticket → r44 → Get_Room_Failed → policy, then size it before acting. Same path fits the next attribution question.
Cases 模块如何把一条用户投诉变成通话遥测 —— 以及在工单 1376296 上,如何发现并量化了一类根本不是质量问题的「no-ring」。
Cases 是 Lighthouse 的「自下而上」那一半:从一个真实信号出发 —— 一张客服工单或一次低星评分 —— 反向走到能解释它的遥测。它把投诉对应到具体用户,直接从 footprint 拉出该用户近期的 VoIP 通话,让你按客户端真实体验逐通查看。
它替你隐藏的难点是「身份」:客服工单用支持系统的 user-id,footprint 用手机号(MSISDN),评分又是另一套 id。Cases 把它们桥接起来(文本里的手机号 → 作者历史 → room-id),让「这条投诉」和「这些通话」对得上。
访问 /cases/<ticketId>。Cases 拉取客服工单、提取手机号、对应到 footprint 用户。你会看到投诉正文、分类,以及首个机器判定。
一眼框定工单:通话量、接通率,以及这些通话聚集在什么上 —— 机型、客户端版本、专线 relay、ISP —— 外加 Top 结束原因和红旗提示(几乎不接通 / 单一版本 / send-mute / 重试风暴)。
每通一行:结果、结束原因 rN · 标签、RTT / 丢包、relay、机型,以及用户自己的 ★ 评分。坏值会标红,问题通话一眼可见。
完整逐通详情 —— 音频内部(PLC / 抖动缓冲 / quality head→tail)、视频、传输深度、以及权限(麦克风 / 相机)。每个原因都带标签和指纹证据。「双侧对比」把用户这条腿和对端并排,区分服务端 / relay 问题还是本地网络问题。
对「打不通 / no-ring」的通话,客户端说不清原因 —— 顺着通话的 traceId 进服务端 call-trace:S1 chat-server 起呼 → S2 im-push → S3 FCM/APNs。哪一段失败,就点出是谁的锅。
复制或下载完整上下文 —— 遥测、原因标签、标识符 —— 到 Sherlock 调试 agent 里继续查。
读画像条 + 钻一通失败的电话,得到一条干净的证据链:
| 信号 | 取值 | 读作 |
|---|---|---|
| 窗口内通话 | 99 | 整个投诉窗口 |
| 接通 | 3 · 3% | 几乎不接通 |
| 结束原因 r44 | 89 · 90% | 外呼建立失败 |
| r44 空 roomId | 100% | 从没拿到房间 |
| 服务端 S1 | Get_Room_Failed | 服务端拒绝分房 |
这些通话从不响铃,是因为是我们拒绝了房间(一次刻意的风控动作),却落进了喂给 bad-call 率的同一个 no-ring 桶 —— 指标把一次策略动作当成了质量故障。封 ≠ 坏。
全网量化(8/4–8/10):排除风控封控的 no-ring 后,总 bad-call 率视口径最高变动 0.6% —— 对一个被盯得这么细的率来说,这是一次真实、系统性的移动。
交叉核对:r44 占「从未响铃」约 1.7%,风控封控约占全部 no-ring 的 0.35–0.42% —— 是系统性误计,不是噪声。它确实在移动这个数,而且是不该被移动的那个数。
把风控 Get_Room_Failed 从 no-ring 里单列,让 bad-call 率只计真正的质量故障 —— 而不是我们主动做的策略动作。
8/4–8/10 的量化已做成 widget;把 r44 / Get_Room_Failed 和真实建立失败持续拆开,占比可被监控而非每次重算。
这些用户是被设计性封控的。客服应看到「被策略限制」状态,而不是去追一个并不存在的通话质量 bug。
真正的价值是这套流程:工单 → r44 → Get_Room_Failed → 策略,然后先量化再行动。下一个归因问题照走。