{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-09-28T06:03:00.468Z","headline":"LangSmith 在医疗 AI 中的常见用例：构建可靠性层","description":"LangSmith 可帮助医疗 AI 团队把临床审查转化为可复用的评估器、数据集和发布门禁，从而让 AI 在生产环境中更安全地运行。","url":"https://www.aioga.com/news/cmucxps84003eronicxia0kn2/","mainEntityOfPage":"https://www.aioga.com/news/cmucxps84003eronicxia0kn2/","datePublished":"2026-09-22T17:08:51.000Z","dateModified":"2026-09-22T17:08:51.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://www.langchain.com/blog/reliability-healthcare-ai-langsmith-use-cases","https://aihot.news/items/cmucxps84003eronicxia0kn2"],"canonicalUrl":"https://www.aioga.com/news/cmucxps84003eronicxia0kn2/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：LangSmith 可帮助医疗 AI 团队把临床审查转化为可复用的评估器、数据集和发布门禁，从而让 AI 在生产环境中更安全地运行。 Aioga 将其归入「行业动态」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmucxps84003eronicxia0kn2/","dateCreated":"2026-09-22T17:08:51.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"langchain.com source article","url":"https://www.langchain.com/blog/reliability-healthcare-ai-langsmith-use-cases","datePublished":"2026-09-22T17:08:51.000Z","provider":{"@type":"Organization","name":"langchain.com","url":"https://www.langchain.com/blog/reliability-healthcare-ai-langsmith-use-cases"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.news/items/cmucxps84003eronicxia0kn2","datePublished":"2026-09-22T17:08:51.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.news/items/cmucxps84003eronicxia0kn2"}}],"aggregationSource":"LangChain：Blog（RSS）","originalPublisher":{"name":"langchain.com","url":"https://www.langchain.com/blog/reliability-healthcare-ai-langsmith-use-cases"},"geoDeepAnswer":null,"article":{"id":"cmucxps84003eronicxia0kn2","slug":"cmucxps84003eronicxia0kn2","url":"https://www.aioga.com/news/cmucxps84003eronicxia0kn2/","title":"LangSmith 在医疗 AI 中的常见用例：构建可靠性层","title_en":"","summary":"LangSmith 可帮助医疗 AI 团队把临床审查转化为可复用的评估器、数据集和发布门禁，从而让 AI 在生产环境中更安全地运行。","source":"LangChain：Blog（RSS）","sourceUrl":"https://www.langchain.com/blog/reliability-healthcare-ai-langsmith-use-cases","aiHotUrl":"https://aihot.news/items/cmucxps84003eronicxia0kn2","publishedAt":"2026-09-22T17:08:51.000Z","category":"行业动态","score":58,"selected":false,"articleBody":["In healthcare, the best judges of whether an AI system works correctly are the people whose time it was built to protect. A clinician can quickly tell if a generated note correctly attributes a symptom or a patient was recommended the most appropriate level of care. But they cannot perform this review across thousands of encounters indefinitely. At scale, expert validation becomes the limiting factor.","Healthcare teams have approached this challenge in different ways, but many share one principle: they treat clinical review as infrastructure rather than a recurring operational cost. This blog covers how two organizations, building AI healthcare products in different ways, converge on evaluation practices. Using LangSmith：https://docs.langchain.com/langsmith/observability, they convert expert input into durable assets such as labeled datasets, calibrated evaluators, and automated release gates. The goal is not to eliminate human judgment, but to make its value compound.","Healthcare AI agents operate across various workflows, such as patient care decisions and visit documentation.","Included Health built Dot, an AI guide powered by LangGraph：https://docs.langchain.com/oss/python/langgraph/overview and Deep Agents：https://docs.langchain.com/oss/python/deepagents/overview. It interprets ambiguous member needs, answers coverage and billing questions, routes people to appropriate care, and detects emergencies. A question about whether a scan is covered may reveal, several turns later, that the member actually needs to speak with a primary care physician. Evaluating Dot means both checking its accuracy and whether it used the member's full context to recommend a safe and appropriate next step.","Abridge transforms patient-clinician conversations into clinical notes. With a patient's consent, a physician records the visit, and Abridge converts the conversation into a note that becomes part of the longitudinal health record and supports billing. In this setting, attribution is critical. If a patient's observation is presented as a physician's conclusion, a symptom can become a billable diagnosis. Hallucinations create a different risk: a medication or dosage that was never prescribed can enter the record. A trustworthy note must preserve who said what, capture what matters clinically, and introduce nothing the conversation does not support.","These systems fail in different ways, but both teams use LangSmith to address the same constraint: accuracy is determined by someone outside the engineering team, and that person's time is often the scarcest resource in the system.","That makes expert review both indispensable and a potential bottleneck. As the Abridge team puts it: “trust is earned in drops and lost in buckets.” The goal is to make each expert judgment reusable across future tests, releases, and iterations.","Before clinical judgment is encoded into automated evaluations, teams must first define exactly what experts are evaluating. Three properties make that judgment difficult to scale.","Ground truth is rarely singular. A clinical note does not have one canonical form . What belongs varies by specialty, encounter, and clinician, and reasonable experts can disagree. Reference notes are useful, but treating a single reference as the only correct answer can penalize valid variation while still missing clinically important errors.","Correct inaction matters. Healthcare teams must evaluate whether a system acted correctly and if it recognized when not to act. For instance, Included Health's reviewers check that emergency guardrails trigger when appropriate and remain inactive in benign cases. For example, Abridge tests whether its agent stays within its boundaries and selects the tools a clinician would expect.","Reviewer expertise is part of the specification. Abridge determines upfront whether an evaluation requires a board-certified physician or a particular specialist. A judge is only as good as the judgments that calibrated it.","None of this makes clinical judgment impossible to automate. It simply defines the requirements for doing so responsibly: tolerance for valid variation, attention to what did not happen, and the right expertise behind every label.","Once teams have defined the judgment they need to preserve, they can begin converting expertise into infrastructure. Abridge converts clinician input into labeled datasets and calibrated judges that keep working after an individual review ends.","Abridge begins with known failure modes from clinician and user feedback. The team ranks them by prevalence and severity, groups them into categories such as accuracy, compliance, style, and completeness, and builds a separate judge for each. Rather than produce one general quality score, the evaluators test for specific ways a note can fail.","The time savings come from automating what happens after they've provided their judgment . Previously, a clinician wrote an annotation guide and labeled encounters, then someone manually adjusted the judge’s prompt until its scores matched those labels. Abridge now feeds the same guide and examples into an automated prompt optimization framework that generates the judge.","Because no single reference can capture every valid note, Abridge layers two approaches with complementary strengths:","Together, they balance breadth and precision: one provides scalable coverage across encounters, while the other captures the specialty-specific nuance that clinical review demands.","An optimized judge still must be validated against clinician annotations. LangSmith’s Align Evaluator：https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback gives teams an interface for comparing the two and investigating disagreements. Abridge separately asks annotators to explain their decisions, even when the output is correct. Those explanations help resolve inconsistencies and confirm the labels reflect careful review.","The result is an evaluation system that can be inspected, recalibrated, and reused. By turning individual judgments into durable evaluators, teams reserve scarce clinical expertise for the cases where it adds the most value. Production review supplies the new cases and feedback that keep those standards current.","Calibrated judges apply judgment the team has already captured. Production review supplies the next round of that judgment, revealing how the system behaves in real conversations and generating evidence for what to fix.","Included Health shows how that new evidence enters the loop. Conversations go into a LangSmith annotation queue：https://docs.langchain.com/langsmith/annotation-queues, where clinical reviewers assess whether Dot directed the member to the right care setting, whether its emergency guardrails behaved appropriately, and whether the case requires follow-up.","Those decisions become structured labels that are exported to Included Health's data warehouse, where the data science team uses them to build operational dashboards. They also feed back into the skill definitions that govern how Dot navigates members. Each review is spent once and used three times.","The result is a powerful feedback loop. Clinical judgment becomes data, the data guides product changes, and those changes are tested against the same standards before the next release. Every review contributes to both the case at hand and the system’s future behavior.","The feedback loop pays off at release time. Instead of evaluating every candidate change from scratch, teams can test it against evidence they have already captured.","At Abridge, a model change moves through progressively more realistic stages: offline evaluations, backtesting against historical encounters, a limited A/B test, full release, and continuous production monitoring. Each stage adds a different kind of evidence.","The A/B test is the most unusual step in this process. Some of Abridge’s partners agree to be among the first 10 to 15% of customers included in a silent rollout. This lets Abridge observe signals automated judges cannot provide: whether clinicians edit the generated notes, how they rate them, and what qualitative feedback they share. Offline evaluations establish whether a change is ready for limited exposure; production behavior determines whether the rollout should expand. That process reduced Abridge’s release cycle from one or two months to a matter of days.","Included Health applied the same principle to an architectural change. Moving Dot’s supergraph to Deep Agents affected four product teams, all wary of breaking changes. The team ran its existing multi-turn simulation suite, confirmed that performance held, and completed the migration in under two weeks without significant regressions.","In both cases, release confidence became cumulative. Rather than re-establish trust with every change, teams could build on evidence they had already collected.","A faster release cycle matters only if the system performs reliably once it reaches real users. Included Health measures performance across three dimensions: adoption, routing quality, and safety.","Each metric answers a different question: Will members use the product? Does it direct them to appropriate care? Does it recognize situations that require urgent attention? Looking at them together gives the team a more complete picture of production reliability.","Following Dot's launch, Included Health reports a 75% lift in chat engagement. Among the graded conversations, clinician agreement with Dot’s care recommendations remains above the team’s 95% target, and clinical audits show that Dot identifies more than 99% of high-risk situations.","At Abridge, labeled encounters calibrate judges that run against future releases; at Included Health, clinical labels outlive the conversation that produced them. The artifacts still require review and recalibration, but the expert judgment behind them is no longer consumed by a single decision.","The same artifacts that make clinical judgment reusable—encounter traces, conversation histories, and clinician annotations—can also contain protected health information. Once teams begin storing and reusing them, security and deployment architecture become part of the evaluation design.","Abridge treats self-hosting, access controls, and auditability as requirements for its evaluation infrastructure. They also remove identifying information from conversation data before using it for learning. These are not controls to add after the evaluation pipeline is built; they shape what data can enter it in the first place.","LangSmith supports managed cloud：https://docs.langchain.com/langsmith/cloud, bring-your-own-cloud：https://docs.langchain.com/langsmith/byoc, and self-hosted：https://docs.langchain.com/langsmith/self-hosted deployments. Teams must decide where evaluation data will be stored, who can access it, which audit and retention controls apply, and how traces containing PHI will be handled.","Trust builds slowly in healthcare AI. It grows with every encounter handled correctly, every guardrail, and every regression caught before it reaches users. Yet one change that escapes those checks can undo it.","Healthcare teams move fast by ensuring each careful review continues working long after the review itself is complete.","For the complete customer stories, watch Building Clinical AI Agents with LangGraph: Abridge's Eval Stack for High-Stakes Healthcare：https://www.youtube.com/watch?v=mxweSHetuN8&t=356s and read How Included Health Built Federated Agents for Healthcare Navigation with Deep Agents and LangGraph：https://www.langchain.com/blog/how-included-health-built-federated-agents-for-healthcare-navigation-with-deep-agents-and-langgraph.","LangSmith, our agent engineering platform, helps developers debug every agent decision, eval changes, and deploy in one click."],"articleImages":[],"mediaStatus":"none","articleBodyZh":["在医疗保健中，判断一个 AI 系统是否运行正确的最佳人选是那些其设计目的是为了保护其时间的人。临床医生可以快速判断生成的记录是否正确标注了症状，或者患者是否被推荐了最合适的护理等级。但他们无法无限期地在成千上万个病例中执行这种审核。在大规模情况下，专家验证成为限制因素。","医疗团队以不同方式应对这一挑战，但许多团队共享一个原则：将临床审查视为基础设施，而非经常性的运营成本。本文博客介绍了两个以不同方式构建 AI 医疗产品的组织如何在评估实践上趋同。通过使用 LangSmith：https://docs.langchain.com/langsmith/observability，他们将专家输入转化为持久资产，例如标注数据集、校准评估器和自动发布门。目标不是消除人工判断，而是使其价值可以复利增长。","医疗 AI 代理在各种工作流程中运行，例如患者护理决策和就诊记录。","Included Health 构建了 Dot，这是一个由 LangGraph：https://docs.langchain.com/oss/python/langgraph/overview 和 Deep Agents：https://docs.langchain.com/oss/python/deepagents/overview 提供支持的 AI 指导。它可以解读模糊的会员需求，回答保险和计费相关问题，将人引导至合适的护理，并检测紧急情况。关于扫描是否涵盖的问题可能在数次交互后显示，会员实际上需要与初级保健医生交谈。评估 Dot 意味着既要检查其准确性，也要确认它是否利用会员的完整背景来推荐安全、适当的下一步措施。","Abridge 将患者和临床医生的对话转换为临床记录。在获得患者同意的情况下，医生会记录就诊内容，而 Abridge 则将对话转换为一份记录，成为纵向健康记录的一部分，并支持计费。在这种情况下，归属至关重要。如果将患者的观察呈现为医生的结论，症状可能成为可计费的诊断。幻觉则带来了不同的风险：从未开过的药物或剂量可能进入记录。一份值得信赖的记录必须保留是谁说了什么，捕捉临床上重要的内容，并且不引入对话中没有支持的内容。","这些系统以不同方式出错，但两支团队都使用 LangSmith 来应对相同的约束：准确性由工程团队之外的人来确定，而该人的时间往往是系统中最稀缺的资源。","这使得专家审查既不可或缺，又可能成为瓶颈。正如 Abridge 团队所说：“信任是逐滴赢得，却一桶桶失去。”目标是让每一次专家判断在未来的测试、发布和迭代中可重复使用。","在临床判断被纳入自动评估之前，团队必须首先明确定义专家究竟在评估什么。三个属性使得这种判断难以规模化推广。","真实情况很少是单一的。临床记录没有一个标准的唯一形式。内容的归属因专业、就诊情况和临床医生而异，合理的专家可能会有不同意见。参考记录很有用，但将单一参考作为唯一正确答案可能会惩罚有效的变异，同时仍可能漏掉临床上重要的错误。","正确的不作为也很重要。医疗团队必须评估系统的行为是否正确，以及它是否识别出何时不该采取行动。例如，Included Health 的审查员会检查紧急防护措施在适当时触发，并在良性情况下保持不活跃。例如，Abridge 会测试其代理是否保持在自身界限内，并选择临床医生预期的工具。","审查员的专业知识是规格的一部分。Abridge 会事先确定评估是否需要具有执照的主治医生或特定专家。一个评审员的好坏取决于用来校准的判断质量。","这些都不使临床判断无法自动化。它只是定义了负责任地进行自动化的要求：对有效差异的容忍、关注未发生的情况，以及每个标签背后的正确专业知识。","一旦团队定义了需要保留的判断，他们就可以开始将专业知识转化为基础设施。Abridge 将临床医生的输入转换为带标签的数据集和经过校准的评审员，这些评审员在个人审查结束后仍继续工作。","Abridge 从临床医生和用户反馈中已知的失败模式开始。团队按普遍性和严重性对这些模式进行排名，将其分组为准确性、合规性、风格和完整性等类别，并为每个类别构建一个独立的评审员。评估者不是产生一个通用的质量分数，而是针对笔记可能出错的具体方式进行测试。","时间节省来自于自动化处理他们提供判断后的步骤。之前，临床医生编写注释指南并标注病例，然后有人手动调整评审员提示，直到其分数与这些标签匹配。现在，Abridge 将相同的指南和示例输入到自动提示优化框架中，从而生成评审员。","由于没有单一参考可以涵盖所有有效笔记，Abridge 结合了两种具有互补优势的方法：","两者共同在广度和精准度之间达到平衡：一种方法提供覆盖所有病例的可扩展性，另一种方法捕捉临床审查所需的专业细微差别。","优化后的评审员仍必须通过临床医生注释进行验证。LangSmith 的 Align Evaluator：https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback 为团队提供了一个界面，用于比较两者并调查分歧。Abridge 还单独要求注释者解释他们的决定，即使输出是正确的。这些解释有助于解决不一致并确认标签反映了认真审查。","最终结果是一个可以检查、重新校准和重复使用的评估系统。通过将个人判断转化为持久的评审员，团队将稀缺的临床专业知识保留用于最有价值的病例。生产审查提供新的病例和反馈，使这些标准保持最新。","经过校准的评审者会应用团队已经捕捉到的判断。生产审查提供下一轮判断，揭示系统在真实对话中的表现，并生成用于修正的证据。","Included Health 展示了这些新证据如何进入循环。对话进入 LangSmith 注释队列：https://docs.langchain.com/langsmith/annotation-queues，由临床评审员评估 Dot 是否将成员引导至正确的护理环境，其紧急防护措施是否运作得当，以及该案例是否需要后续处理。","这些决策变成结构化标签，导出到 Included Health 的数据仓库，数据科学团队使用这些数据构建运营仪表板。它们也反馈到控制 Dot 导航成员的技能定义中。每次审查使用一次，却被三次利用。","结果是一个强大的反馈循环。临床判断变成数据，数据指导产品变更，这些变更在下次发布前根据相同标准进行测试。每次审查都对当前案例和系统的未来行为有所贡献。","在发布时，这个反馈循环会带来回报。团队不必从头评估每个候选变更，而可以用已捕获的证据进行测试。","在 Abridge，模型变更通过逐步更现实的阶段进行：离线评估、对历史就诊记录的回测、有限的 A/B 测试、全面发布以及持续的生产监控。每个阶段都增加不同类型的证据。","A/B 测试是此过程中最不寻常的步骤。一些 Abridge 的合作伙伴同意成为静默发布中首批 10% 到 15% 的客户。这让 Abridge 能观察自动化评审无法提供的信号：临床医生是否修改生成的记录，他们如何评分，以及他们提供了什么定性反馈。离线评估确定变更是否准备好进行有限曝光；生产行为决定是否扩大发布。该过程将 Abridge 的发布周期从一两个月缩短到数天。","Included Health 将相同的原则应用于架构变更。将 Dot 的 supergraph 移动到 Deep Agents 影响了四个产品团队，这些团队都对破坏性变更保持警惕。团队运行了现有的多回合模拟测试套件，确认性能保持稳定，并在不到两周的时间内完成了迁移，没有出现显著的回归。","在这两种情况下，发布信心都是累积的。团队不必在每次变更时重新建立信任，而可以基于他们已经收集的证据继续构建。","只有当系统在真正用户中表现可靠时，更快的发布周期才有意义。Included Health 从三个维度测量性能：采用率、引导质量和安全性。","每个指标回答不同的问题：成员会使用该产品吗？它能将他们引导到合适的护理吗？它能识别需要紧急关注的情况吗？将这些指标综合来看，可以为团队提供生产可靠性的更完整图景。","在 Dot 发布后，Included Health 报告聊天参与度提升了 75%。在经过评分的对话中，临床医生对 Dot 的护理建议的认可率仍保持在团队设定的 95% 目标以上，临床审核显示 Dot 能识别超过 99% 的高风险情况。","在 Abridge，带标签的就诊记录用于校准针对未来版本的评审；在 Included Health，临床标签的有效性超越了生成它们的对话。这些文档仍需复审和重新校准，但其背后的专家判断不再仅用于单一决策。","使临床判断可重复使用的相同文档——就诊记录痕迹、对话历史和临床标注——也可能包含受保护的健康信息。一旦团队开始存储和重用这些信息，安全性和部署架构就成为评估设计的一部分。","Abridge 将自托管、访问控制和可审计性视为其评估基础设施的要求。他们还会在使用对话数据进行学习之前，移除其中的身份信息。这些不是在评估管道构建完成后再添加的控制措施，而是决定了最初哪些数据可以进入管道的设计因素。","LangSmith 支持托管云：https://docs.langchain.com/langsmith/cloud、自带云：https://docs.langchain.com/langsmith/byoc 和自托管：https://docs.langchain.com/langsmith/self-hosted 部署。团队必须决定评估数据将存储在哪里，谁可以访问，适用哪些审计和保留控制，以及如何处理包含 PHI 的追踪记录。","信任在医疗 AI 中建立得很慢。它随着每一次正确处理的交互、每一个安全防护以及每一次在到达用户之前被发现的回归而增长。然而，一次逃过这些检查的变更就可能将其摧毁。","医疗团队通过确保每一次细致的审查在审查完成后仍能持续有效，从而快速推进工作。","欲了解完整的客户故事，请观看《使用 LangGraph 构建临床 AI 代理：Abridge 的高风险医疗评估堆栈》：https://www.youtube.com/watch?v=mxweSHetuN8&t=356s，并阅读《Included Health 如何使用 Deep Agents 和 LangGraph 构建医疗导航的联合代理》：https://www.langchain.com/blog/how-included-health-built-federated-agents-for-healthcare-navigation-with-deep-agents-and-langgraph。","LangSmith，我们的代理工程平台，帮助开发者调试每一个代理决策、评估变更，并一键部署。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：LangSmith 可帮助医疗 AI 团队把临床审查转化为可复用的评估器、数据集和发布门禁，从而让 AI 在生产环境中更安全地运行。 Aioga 将其归入「行业动态」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：公司与行业类动态需要放在竞争格局、商业化路径、资本信号和监管环境中观察，单条公告不能代表最终结果。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察官方文件、合作落地、收入或用户信号、竞品动作和监管后续。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-09-28T06:19:15.377Z","sourceHash":"0e2d3db1c2c7af8b","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["行业动态","LangChain：Blog（RSS）"],"translations":{"zh-CN":{"title":"LangSmith 在医疗 AI 中的常见用例：构建可靠性层","summary":"LangSmith 可帮助医疗 AI 团队把临床审查转化为可复用的评估器、数据集和发布门禁，从而让 AI 在生产环境中更安全地运行。","category":"行业动态","source":"langchain.com","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangSmith 在医疗 AI 中的常见用例：构建可靠性层 - Aioga AI资讯","description":"LangSmith 可帮助医疗 AI 团队把临床审查转化为可复用的评估器、数据集和发布门禁，从而让 AI 在生产环境中更安全地运行。","url":"https://www.aioga.com/news/cmucxps84003eronicxia0kn2/","articleBody":["在医疗保健中，判断一个 AI 系统是否运行正确的最佳人选是那些其设计目的是为了保护其时间的人。临床医生可以快速判断生成的记录是否正确标注了症状，或者患者是否被推荐了最合适的护理等级。但他们无法无限期地在成千上万个病例中执行这种审核。在大规模情况下，专家验证成为限制因素。","医疗团队以不同方式应对这一挑战，但许多团队共享一个原则：将临床审查视为基础设施，而非经常性的运营成本。本文博客介绍了两个以不同方式构建 AI 医疗产品的组织如何在评估实践上趋同。通过使用 LangSmith：https://docs.langchain.com/langsmith/observability，他们将专家输入转化为持久资产，例如标注数据集、校准评估器和自动发布门。目标不是消除人工判断，而是使其价值可以复利增长。","医疗 AI 代理在各种工作流程中运行，例如患者护理决策和就诊记录。","Included Health 构建了 Dot，这是一个由 LangGraph：https://docs.langchain.com/oss/python/langgraph/overview 和 Deep Agents：https://docs.langchain.com/oss/python/deepagents/overview 提供支持的 AI 指导。它可以解读模糊的会员需求，回答保险和计费相关问题，将人引导至合适的护理，并检测紧急情况。关于扫描是否涵盖的问题可能在数次交互后显示，会员实际上需要与初级保健医生交谈。评估 Dot 意味着既要检查其准确性，也要确认它是否利用会员的完整背景来推荐安全、适当的下一步措施。","Abridge 将患者和临床医生的对话转换为临床记录。在获得患者同意的情况下，医生会记录就诊内容，而 Abridge 则将对话转换为一份记录，成为纵向健康记录的一部分，并支持计费。在这种情况下，归属至关重要。如果将患者的观察呈现为医生的结论，症状可能成为可计费的诊断。幻觉则带来了不同的风险：从未开过的药物或剂量可能进入记录。一份值得信赖的记录必须保留是谁说了什么，捕捉临床上重要的内容，并且不引入对话中没有支持的内容。","这些系统以不同方式出错，但两支团队都使用 LangSmith 来应对相同的约束：准确性由工程团队之外的人来确定，而该人的时间往往是系统中最稀缺的资源。","这使得专家审查既不可或缺，又可能成为瓶颈。正如 Abridge 团队所说：“信任是逐滴赢得，却一桶桶失去。”目标是让每一次专家判断在未来的测试、发布和迭代中可重复使用。","在临床判断被纳入自动评估之前，团队必须首先明确定义专家究竟在评估什么。三个属性使得这种判断难以规模化推广。","真实情况很少是单一的。临床记录没有一个标准的唯一形式。内容的归属因专业、就诊情况和临床医生而异，合理的专家可能会有不同意见。参考记录很有用，但将单一参考作为唯一正确答案可能会惩罚有效的变异，同时仍可能漏掉临床上重要的错误。","正确的不作为也很重要。医疗团队必须评估系统的行为是否正确，以及它是否识别出何时不该采取行动。例如，Included Health 的审查员会检查紧急防护措施在适当时触发，并在良性情况下保持不活跃。例如，Abridge 会测试其代理是否保持在自身界限内，并选择临床医生预期的工具。","审查员的专业知识是规格的一部分。Abridge 会事先确定评估是否需要具有执照的主治医生或特定专家。一个评审员的好坏取决于用来校准的判断质量。","这些都不使临床判断无法自动化。它只是定义了负责任地进行自动化的要求：对有效差异的容忍、关注未发生的情况，以及每个标签背后的正确专业知识。","一旦团队定义了需要保留的判断，他们就可以开始将专业知识转化为基础设施。Abridge 将临床医生的输入转换为带标签的数据集和经过校准的评审员，这些评审员在个人审查结束后仍继续工作。","Abridge 从临床医生和用户反馈中已知的失败模式开始。团队按普遍性和严重性对这些模式进行排名，将其分组为准确性、合规性、风格和完整性等类别，并为每个类别构建一个独立的评审员。评估者不是产生一个通用的质量分数，而是针对笔记可能出错的具体方式进行测试。","时间节省来自于自动化处理他们提供判断后的步骤。之前，临床医生编写注释指南并标注病例，然后有人手动调整评审员提示，直到其分数与这些标签匹配。现在，Abridge 将相同的指南和示例输入到自动提示优化框架中，从而生成评审员。","由于没有单一参考可以涵盖所有有效笔记，Abridge 结合了两种具有互补优势的方法：","两者共同在广度和精准度之间达到平衡：一种方法提供覆盖所有病例的可扩展性，另一种方法捕捉临床审查所需的专业细微差别。","优化后的评审员仍必须通过临床医生注释进行验证。LangSmith 的 Align Evaluator：https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback 为团队提供了一个界面，用于比较两者并调查分歧。Abridge 还单独要求注释者解释他们的决定，即使输出是正确的。这些解释有助于解决不一致并确认标签反映了认真审查。","最终结果是一个可以检查、重新校准和重复使用的评估系统。通过将个人判断转化为持久的评审员，团队将稀缺的临床专业知识保留用于最有价值的病例。生产审查提供新的病例和反馈，使这些标准保持最新。","经过校准的评审者会应用团队已经捕捉到的判断。生产审查提供下一轮判断，揭示系统在真实对话中的表现，并生成用于修正的证据。","Included Health 展示了这些新证据如何进入循环。对话进入 LangSmith 注释队列：https://docs.langchain.com/langsmith/annotation-queues，由临床评审员评估 Dot 是否将成员引导至正确的护理环境，其紧急防护措施是否运作得当，以及该案例是否需要后续处理。","这些决策变成结构化标签，导出到 Included Health 的数据仓库，数据科学团队使用这些数据构建运营仪表板。它们也反馈到控制 Dot 导航成员的技能定义中。每次审查使用一次，却被三次利用。","结果是一个强大的反馈循环。临床判断变成数据，数据指导产品变更，这些变更在下次发布前根据相同标准进行测试。每次审查都对当前案例和系统的未来行为有所贡献。","在发布时，这个反馈循环会带来回报。团队不必从头评估每个候选变更，而可以用已捕获的证据进行测试。","在 Abridge，模型变更通过逐步更现实的阶段进行：离线评估、对历史就诊记录的回测、有限的 A/B 测试、全面发布以及持续的生产监控。每个阶段都增加不同类型的证据。","A/B 测试是此过程中最不寻常的步骤。一些 Abridge 的合作伙伴同意成为静默发布中首批 10% 到 15% 的客户。这让 Abridge 能观察自动化评审无法提供的信号：临床医生是否修改生成的记录，他们如何评分，以及他们提供了什么定性反馈。离线评估确定变更是否准备好进行有限曝光；生产行为决定是否扩大发布。该过程将 Abridge 的发布周期从一两个月缩短到数天。","Included Health 将相同的原则应用于架构变更。将 Dot 的 supergraph 移动到 Deep Agents 影响了四个产品团队，这些团队都对破坏性变更保持警惕。团队运行了现有的多回合模拟测试套件，确认性能保持稳定，并在不到两周的时间内完成了迁移，没有出现显著的回归。","在这两种情况下，发布信心都是累积的。团队不必在每次变更时重新建立信任，而可以基于他们已经收集的证据继续构建。","只有当系统在真正用户中表现可靠时，更快的发布周期才有意义。Included Health 从三个维度测量性能：采用率、引导质量和安全性。","每个指标回答不同的问题：成员会使用该产品吗？它能将他们引导到合适的护理吗？它能识别需要紧急关注的情况吗？将这些指标综合来看，可以为团队提供生产可靠性的更完整图景。","在 Dot 发布后，Included Health 报告聊天参与度提升了 75%。在经过评分的对话中，临床医生对 Dot 的护理建议的认可率仍保持在团队设定的 95% 目标以上，临床审核显示 Dot 能识别超过 99% 的高风险情况。","在 Abridge，带标签的就诊记录用于校准针对未来版本的评审；在 Included Health，临床标签的有效性超越了生成它们的对话。这些文档仍需复审和重新校准，但其背后的专家判断不再仅用于单一决策。","使临床判断可重复使用的相同文档——就诊记录痕迹、对话历史和临床标注——也可能包含受保护的健康信息。一旦团队开始存储和重用这些信息，安全性和部署架构就成为评估设计的一部分。","Abridge 将自托管、访问控制和可审计性视为其评估基础设施的要求。他们还会在使用对话数据进行学习之前，移除其中的身份信息。这些不是在评估管道构建完成后再添加的控制措施，而是决定了最初哪些数据可以进入管道的设计因素。","LangSmith 支持托管云：https://docs.langchain.com/langsmith/cloud、自带云：https://docs.langchain.com/langsmith/byoc 和自托管：https://docs.langchain.com/langsmith/self-hosted 部署。团队必须决定评估数据将存储在哪里，谁可以访问，适用哪些审计和保留控制，以及如何处理包含 PHI 的追踪记录。","信任在医疗 AI 中建立得很慢。它随着每一次正确处理的交互、每一个安全防护以及每一次在到达用户之前被发现的回归而增长。然而，一次逃过这些检查的变更就可能将其摧毁。","医疗团队通过确保每一次细致的审查在审查完成后仍能持续有效，从而快速推进工作。","欲了解完整的客户故事，请观看《使用 LangGraph 构建临床 AI 代理：Abridge 的高风险医疗评估堆栈》：https://www.youtube.com/watch?v=mxweSHetuN8&t=356s，并阅读《Included Health 如何使用 Deep Agents 和 LangGraph 构建医疗导航的联合代理》：https://www.langchain.com/blog/how-included-health-built-federated-agents-for-healthcare-navigation-with-deep-agents-and-langgraph。","LangSmith，我们的代理工程平台，帮助开发者调试每一个代理决策、评估变更，并一键部署。"]},"en":{"title":"Common Use Cases of LangSmith in Healthcare AI: Building a Reliability Layer","summary":"LangSmith can help healthcare AI teams turn clinical reviews into reusable evaluators, datasets, and release gates, thereby allowing AI to operate more safely in production environments.","category":"Industry","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Common Use Cases of LangSmith in Healthcare AI: Building a Reliability Layer - Aioga AI News","description":"LangSmith can help healthcare AI teams turn clinical reviews into reusable evaluators, datasets, and release gates, thereby allowing AI to operate more safely in production environ...","url":"https://www.aioga.com/en/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:05:38.868Z"},"ja":{"title":"LangSmith の医療 AI における一般的なユースケース：信頼性レイヤーの構築","summary":"LangSmith は、医療 AI チームが臨床レビューを再利用可能な評価器、データセット、および公開ゲートに変換するのを支援し、それによって AI が本番環境でより安全に動作できるようにします。","category":"業界動向","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangSmith の医療 AI における一般的なユースケース：信頼性レイヤーの構築 - Aioga AIニュース","description":"LangSmith は、医療 AI チームが臨床レビューを再利用可能な評価器、データセット、および公開ゲートに変換するのを支援し、それによって AI が本番環境でより安全に動作できるようにします。","url":"https://www.aioga.com/ja/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:05:44.664Z"},"ko":{"title":"LangSmith의 의료 AI에서의 일반적인 사용 사례: 신뢰성 계층 구축","summary":"LangSmith는 의료 AI 팀이 임상 검토를 재사용 가능한 평가자, 데이터 세트 및 배포 접근 제어로 전환하여 AI가 생산 환경에서 더 안전하게 실행되도록 도울 수 있습니다.","category":"업계 동향","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangSmith의 의료 AI에서의 일반적인 사용 사례: 신뢰성 계층 구축 - Aioga AI 뉴스","description":"LangSmith는 의료 AI 팀이 임상 검토를 재사용 가능한 평가자, 데이터 세트 및 배포 접근 제어로 전환하여 AI가 생산 환경에서 더 안전하게 실행되도록 도울 수 있습니다.","url":"https://www.aioga.com/ko/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:06:27.825Z"},"es":{"title":"Casos de uso comunes de LangSmith en la IA médica: construir capas de confiabilidad","summary":"LangSmith puede ayudar a los equipos de IA médica a convertir la revisión clínica en evaluadores, conjuntos de datos y puertas de publicación reutilizables, lo que permite que la IA funcione de manera más segura en entornos de producción.","category":"Industria","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Casos de uso comunes de LangSmith en la IA médica: construir capas de confiabilidad - Aioga Noticias de IA","description":"LangSmith puede ayudar a los equipos de IA médica a convertir la revisión clínica en evaluadores, conjuntos de datos y puertas de publicación reutilizables, lo que permite que la I...","url":"https://www.aioga.com/es/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:06:38.370Z"},"fr":{"title":"Cas d'utilisation courants de LangSmith dans l'IA médicale : construction d'une couche de fiabilité","summary":"LangSmith peut aider les équipes d'IA médicale à transformer les examens cliniques en évaluateurs réutilisables, ensembles de données et contrôles de publication, permettant ainsi à l'IA de fonctionner plus en sécurité en environnement de production.","category":"Industrie","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Cas d'utilisation courants de LangSmith dans l'IA médicale : construction d'une couche de fiabilité - Aioga Actualités IA","description":"LangSmith peut aider les équipes d'IA médicale à transformer les examens cliniques en évaluateurs réutilisables, ensembles de données et contrôles de publication, permettant ainsi...","url":"https://www.aioga.com/fr/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:07:18.320Z"},"de":{"title":"Häufige Anwendungsfälle von LangSmith im Bereich der medizinischen KI: Aufbau einer Zuverlässigkeitsebene","summary":"LangSmith kann medizinischen KI-Teams helfen, klinische Überprüfungen in wiederverwendbare Evaluatoren, Datensätze und Veröffentlichungszugänge umzuwandeln, wodurch KI sicherer in Produktionsumgebungen betrieben werden kann.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Häufige Anwendungsfälle von LangSmith im Bereich der medizinischen KI: Aufbau einer Zuverlässigkeitsebene - Aioga KI-News","description":"LangSmith kann medizinischen KI-Teams helfen, klinische Überprüfungen in wiederverwendbare Evaluatoren, Datensätze und Veröffentlichungszugänge umzuwandeln, wodurch KI sicherer in...","url":"https://www.aioga.com/de/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:07:18.300Z"},"pt-BR":{"title":"Casos de uso comuns do LangSmith em IA médica: Construindo camadas de confiabilidade","summary":"LangSmith pode ajudar equipes de IA na área de saúde a transformar revisões clínicas em avaliadores reutilizáveis, conjuntos de dados e controles de publicação, permitindo que a IA opere de forma mais segura em ambientes de produção.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Casos de uso comuns do LangSmith em IA médica: Construindo camadas de confiabilidade - Aioga Notícias de IA","description":"LangSmith pode ajudar equipes de IA na área de saúde a transformar revisões clínicas em avaliadores reutilizáveis, conjuntos de dados e controles de publicação, permitindo que a IA...","url":"https://www.aioga.com/pt-BR/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:07:58.013Z"},"ru":{"title":"Обычные случаи использования LangSmith в медицинском ИИ: создание слоя надежности","summary":"LangSmith может помочь командам медицинского ИИ преобразовывать клинические обзоры в повторно используемые оценщики, наборы данных и системы управления публикацией, чтобы ИИ работал в производственной среде более безопасно.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Обычные случаи использования LangSmith в медицинском ИИ: создание слоя надежности - Aioga Новости ИИ","description":"LangSmith может помочь командам медицинского ИИ преобразовывать клинические обзоры в повторно используемые оценщики, наборы данных и системы управления публикацией, чтобы ИИ работа...","url":"https://www.aioga.com/ru/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:08:00.558Z"},"ar":{"title":"الاستخدامات الشائعة لـ LangSmith في الذكاء الاصطناعي الطبي: بناء طبقة الموثوقية","summary":"يمكن لـ LangSmith مساعدة فرق الذكاء الاصطناعي الطبي على تحويل المراجعات السريرية إلى مقيمات ومجموعات بيانات وأذونات نشر قابلة لإعادة الاستخدام، مما يسمح للذكاء الاصطناعي بالعمل بأمان أكبر في بيئة الإنتاج.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"الاستخدامات الشائعة لـ LangSmith في الذكاء الاصطناعي الطبي: بناء طبقة الموثوقية - Aioga أخبار الذكاء الاصطناعي","description":"يمكن لـ LangSmith مساعدة فرق الذكاء الاصطناعي الطبي على تحويل المراجعات السريرية إلى مقيمات ومجموعات بيانات وأذونات نشر قابلة لإعادة الاستخدام، مما يسمح للذكاء الاصطناعي بالعمل بأم...","url":"https://www.aioga.com/ar/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:08:42.271Z"},"hi":{"title":"LangSmith चिकित्सा AI में सामान्य उपयोग के मामले: विश्वसनीयता परत बनाना","summary":"LangSmith चिकित्सा AI टीमों को नैदानिक समीक्षा को पुन: प्रयोज्य मूल्यांकनकर्ता, डेटा सेट और प्रकाशन पहुंच नियंत्रण में बदलने में मदद कर सकता है, जिससे AI उत्पादन वातावरण में अधिक सुरक्षित रूप से चल सके।","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangSmith चिकित्सा AI में सामान्य उपयोग के मामले: विश्वसनीयता परत बनाना - Aioga AI समाचार","description":"LangSmith चिकित्सा AI टीमों को नैदानिक समीक्षा को पुन: प्रयोज्य मूल्यांकनकर्ता, डेटा सेट और प्रकाशन पहुंच नियंत्रण में बदलने में मदद कर सकता है, जिससे AI उत्पादन वातावरण में अधिक स...","url":"https://www.aioga.com/hi/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:08:42.531Z"},"it":{"title":"Esempi comuni di utilizzo di LangSmith nell'IA medica: costruire un livello di affidabilità","summary":"LangSmith può aiutare i team di AI medica a trasformare la revisione clinica in valutatori riutilizzabili, set di dati e controlli di pubblicazione, consentendo così all'AI di funzionare in modo più sicuro in ambienti di produzione.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Esempi comuni di utilizzo di LangSmith nell'IA medica: costruire un livello di affidabilità - Aioga Notizie IA","description":"LangSmith può aiutare i team di AI medica a trasformare la revisione clinica in valutatori riutilizzabili, set di dati e controlli di pubblicazione, consentendo così all'AI di funz...","url":"https://www.aioga.com/it/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:09:35.183Z"},"nl":{"title":"Veelvoorkomende use-cases van LangSmith in medische AI: het opbouwen van een betrouwbaarheidslaag","summary":"LangSmith kan medische AI-teams helpen klinische beoordelingen om te zetten in herbruikbare evaluators, datasets en release-poorten, zodat AI veiliger in productieomgevingen kan opereren.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Veelvoorkomende use-cases van LangSmith in medische AI: het opbouwen van een betrouwbaarheidslaag - Aioga AI-nieuws","description":"LangSmith kan medische AI-teams helpen klinische beoordelingen om te zetten in herbruikbare evaluators, datasets en release-poorten, zodat AI veiliger in productieomgevingen kan op...","url":"https://www.aioga.com/nl/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:09:23.573Z"},"tr":{"title":"LangSmith'in tıbbi yapay zekadaki yaygın kullanım senaryoları: Güvenilirlik katmanı oluşturma","summary":"LangSmith, sağlık AI ekiplerinin klinik incelemeleri yeniden kullanılabilir değerlendirme araçlarına, veri kümelerine ve yayın erişim kontrollerine dönüştürmesine yardımcı olarak AI'nın üretim ortamında daha güvenli çalışmasını sağlar.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"LangSmith'in tıbbi yapay zekadaki yaygın kullanım senaryoları: Güvenilirlik katmanı oluşturma - Aioga AI Haberleri","description":"LangSmith, sağlık AI ekiplerinin klinik incelemeleri yeniden kullanılabilir değerlendirme araçlarına, veri kümelerine ve yayın erişim kontrollerine dönüştürmesine yardımcı olarak A...","url":"https://www.aioga.com/tr/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:10:18.573Z"},"vi":{"title":"Các trường hợp sử dụng phổ biến của LangSmith trong AI y tế: Xây dựng lớp độ tin cậy","summary":"LangSmith có thể giúp các nhóm AI y tế biến các đánh giá lâm sàng thành các bộ đánh giá, bộ dữ liệu và kiểm soát phát hành có thể tái sử dụng, từ đó giúp AI hoạt động an toàn hơn trong môi trường sản xuất.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Các trường hợp sử dụng phổ biến của LangSmith trong AI y tế: Xây dựng lớp độ tin cậy - Tin tức AI Aioga","description":"LangSmith có thể giúp các nhóm AI y tế biến các đánh giá lâm sàng thành các bộ đánh giá, bộ dữ liệu và kiểm soát phát hành có thể tái sử dụng, từ đó giúp AI hoạt động an toàn hơn t...","url":"https://www.aioga.com/vi/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:10:26.731Z"},"id":{"title":"Contoh penggunaan umum LangSmith dalam AI medis: Membangun lapisan keandalan","summary":"LangSmith dapat membantu tim AI medis mengubah tinjauan klinis menjadi evaluator, kumpulan data, dan kontrol rilis yang dapat digunakan kembali, sehingga AI dapat beroperasi dengan lebih aman di lingkungan produksi.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Contoh penggunaan umum LangSmith dalam AI medis: Membangun lapisan keandalan - Berita AI Aioga","description":"LangSmith dapat membantu tim AI medis mengubah tinjauan klinis menjadi evaluator, kumpulan data, dan kontrol rilis yang dapat digunakan kembali, sehingga AI dapat beroperasi dengan...","url":"https://www.aioga.com/id/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:11:09.712Z"},"th":{"title":"กรณีการใช้งานทั่วไปของ LangSmith ใน AI ทางการแพทย์: การสร้างชั้นความน่าเชื่อถือ","summary":"LangSmith สามารถช่วยทีม AI ทางการแพทย์เปลี่ยนการตรวจสอบทางคลินิกให้เป็นตัวประเมิน ข้อมูลชุด และการควบคุมการเผยแพร่ที่นำกลับมาใช้ใหม่ได้ ทำให้ AI สามารถทำงานในสภาพแวดล้อมการผลิตได้อย่างปลอดภัยมากยิ่งขึ้น","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"กรณีการใช้งานทั่วไปของ LangSmith ใน AI ทางการแพทย์: การสร้างชั้นความน่าเชื่อถือ - ข่าว AI Aioga","description":"LangSmith สามารถช่วยทีม AI ทางการแพทย์เปลี่ยนการตรวจสอบทางคลินิกให้เป็นตัวประเมิน ข้อมูลชุด และการควบคุมการเผยแพร่ที่นำกลับมาใช้ใหม่ได้ ทำให้ AI สามารถทำงานในสภาพแวดล้อมการผลิตได้อ...","url":"https://www.aioga.com/th/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:11:15.853Z"},"pl":{"title":"Typowe zastosowania LangSmith w medycznej sztucznej inteligencji: budowanie warstwy niezawodności","summary":"LangSmith może pomóc zespołom medycznej sztucznej inteligencji przekształcić przeglądy kliniczne w ponownie wykorzystywalne ewaluatory, zbiory danych i bramki publikacji, dzięki czemu AI może działać bezpieczniej w środowisku produkcyjnym.","category":"行业动态","source":"LangChain：Blog（RSS）","aggregationSource":"LangChain：Blog（RSS）","pageTitle":"Typowe zastosowania LangSmith w medycznej sztucznej inteligencji: budowanie warstwy niezawodności - Aioga Wiadomości AI","description":"LangSmith może pomóc zespołom medycznej sztucznej inteligencji przekształcić przeglądy kliniczne w ponownie wykorzystywalne ewaluatory, zbiory danych i bramki publikacji, dzięki cz...","url":"https://www.aioga.com/pl/news/cmucxps84003eronicxia0kn2/","contentTranslated":true,"sourceHash":"df2f1fb04331f6be","translatedAt":"2026-09-22T18:12:07.605Z"}},"evidenceTier":"verified-news","reviewStatus":"automated-ingest","indexable":true,"editorialCover":"/page-visuals/topic-timeline.png"}}