Original Note

Boundary-Focused Evaluation for Chinese Clinical Diagnosis Normalization

  • papers
  • Original Note
  • Updated: 2026-08-31T10:06:42+08:00
Source Collection
papers
Source Path
papers/CNBEN/Plan&ToRead/outline.md
Type
Original Note
Updated At
2026-08-31T10:06:42+08:00

Boundary-Focused Evaluation for Chinese Clinical Diagnosis Normalization

Subtitle

Strong retrieval and reranking models may rank medically related Chinese diagnosis candidates highly, but clinical normalization requires equivalence rather than topical relevance. This project builds a boundary-focused evaluation protocol and locked Chinese clinical boundary set, then tests whether equivalence-oriented boundary-aware retrievers or rerankers can reduce semantically related but non-equivalent errors under strong baselines and cross-source evaluation.

Slide 1. Boundary-Focused Evaluation for Chinese Clinical Diagnosis Normalization

Message: 讨论这个课题是否适合作为后续研究方向。

  • Focus: Chinese clinical diagnosis normalization
  • Core tension: topical relevance is not clinical equivalence
  • Goal: decide whether the 0-8 week pilot is worth starting

Visual / evidence: title slide with one-line research positioning.

Slide 2. The Core Problem Is Equivalence, Not Topical Relevance

Message: 临床诊断标准化要求 mention 与 candidate 可映射到同一标准概念,而不是只要医学相关。

  • strong retriever/reranker 可能把医学相关候选排高
  • related-but-not-equivalent candidates 会造成错误标准化
  • semantic boundary cases 是 strong baseline 后最值得检查的剩余错误

Visual / evidence: mention -> top-k candidates -> equivalent vs related-but-not-equivalent 流程图。

Slide 3. This Is Not Just Another Entity Linking Model

Message: 方法模块本身不是最强 novelty,课题重点应放在 strong baseline 后的 boundary failure。

  • retriever + reranker 已经是成熟范式
  • hard negatives、definitions、LLM augmentation 都有近邻工作
  • 弱 baseline 上的提升不能证明问题本身有研究价值
  • 本课题主线:证明并评测 remaining semantic boundary failures

Slide 4. The Project Is Driven by Three Research Questions

Message: 课题成败取决于 boundary failure 是否稳定、可评测、可减少。

  • RQ1: strong baseline 后是否仍存在稳定 boundary errors?
  • RQ2: overall metrics 是否掩盖 boundary subset 的错误?
  • RQ3: equivalence-oriented boundary-aware method 是否能减少 related-but-not-equivalent errors?

Visual / evidence: three-question decision flow: existence -> measurement -> reduction.

Slide 5. The Main Contribution Should Be Evaluation-First

Message: 贡献排序应是 boundary-focused evaluation 优先,而不是包装成全新 BEN 方法。

  • boundary-centered analysis under strong baselines
  • semantic boundary taxonomy and locked boundary test set
  • equivalence-oriented retriever/reranker
  • generalization analysis and negative-result reporting

Visual / evidence: contribution priority pyramid.

Slide 6. Boundary Benchmark Design Makes the Problem Testable

Message: locked boundary set 用来区分模型边界错误和数据/候选库问题。

  • Core types: 近义非同义、同部位不同疾病、修饰/程度不同、语义类型冲突、上下位/粒度不一致
  • Separate model boundary error from candidate missing, gold issue, and terminology conflict
  • Keep development boundary set separate from locked boundary test set

Visual / evidence: compact taxonomy table with one example per type.

Slide 7. The Method Tests Equivalence Discrimination Under Strong Baselines

Message: 方法设计服务于 hypothesis test:模型能否从 relevance ranking 转向 equivalence discrimination。

  • Candidate generation: BM25, BGE-M3, Qwen Embedding
  • Strong baseline: Qwen Embedding + Qwen Reranker
  • Boundary-aware signals: definitions, semantic types, hard negatives
  • Training target: mention-candidate equivalence discrimination

Visual / evidence: pipeline diagram from candidate generation to boundary-aware reranking.

Slide 8. Evaluation Must Separate Overall Performance from Boundary Performance

Message: overall metric 只能说明基础能力,boundary metric 才能检验核心假设。

  • Overall: Accuracy, Recall@k, MRR
  • Boundary-specific: Boundary Accuracy, Boundary MRR, confusion reduction rate
  • Key comparison: standard test set vs locked boundary test set
  • Watch for trade-off: boundary improvement should not destroy top-k recall or overall accuracy

Visual / evidence: two-column metric table: overall metrics vs boundary metrics.

Slide 9. The 0-8 Week Pilot Is a Go / No-Go Gate

Message: 先证伪课题可行性,再决定是否投入完整研究周期。

  • 数据和候选库是否稳定
  • strong baseline 后是否仍有足够 boundary errors
  • 是否能构建 200-400 条 pilot boundary cases
  • locked boundary test set 是否可独立构造
  • minimal method 是否有方向性信号

Visual / evidence: go/no-go checklist with pass / risk / fail columns.

Slide 10. Discussion

Message: 这次组会希望得到课题方向判断,而不是确认完整技术细节。

  • semantic boundary 视角是否值得作为研究主线?
  • 数据、标注和 locked set 是否现实?
  • 公开数据(CBLUE...)可用于初步探索,组内是否能提供新数据用于generalization以及更深入的分析
  • 组内能提供的数据会是什么样的: 有无标注,数据格式,数据规模,数据类型...
  • 如果不成立,是否转向 benchmark/error analysis、候选库覆盖率、拒识或资源建设?

Visual / evidence: decision slide: pursue / revise / pivot.

Slide 11. Ideal Flow

插入Ideal Flow流程图

Evidence-backed relations

Source Note · Same Topic

Evidence-backed relations

Related Summary

切换到中文