VOL. 06 / 编辑型法律 AI 情报本周信号已校对 · 06
← 返回首页
ARTICLE DOSSIERREF / 4E60B803
学术🌐 国际一般2026/08/29

音视频图像理解技术研究进展

SOURCE / Harvey-法律 AI Blog · audio video image understanding

原文

audio video image understanding

完整原文

ProductAnalyze the Full Scope of Legal Evidence, Not Just the DocumentsHarvey now supports audio and video transcription along with enhanced image understanding.by Harvey Team•Aug 14, 2026Legal work has never lived entirely in text — exhibit images, deposition recordings, org charts buried in a merger agreement. Until now, Harvey could read the words around these but not the visuals or recordings themselves, which meant pulling files out to review by hand, or sending audio to a transcription vendor and waiting to re-upload it before analysis could even start.That’s why Harvey now supports audio and video transcription along with enhanced image understanding.Enhanced Image UnderstandingHarvey can now visually interpret images, charts, graphs, and diagrams. Upload a standalone image, or a PDF, Word, or PowerPoint file with visual content, and ask your question. Then, Harvey will spot when visual inspection is needed and analyze the relevant page before answering.Deal teams can read ownership structures straight out of a purchase agreement without redrawing them. Litigators can review patent drawings or trademark specimens directly. In-house teams working through chart-heavy filings or board materials can just ask Harvey what a chart shows, instead of transcribing it into text first.Audio File TranscriptionDrop an audio file into Assistant, or upload it to Vault, and Harvey transcribes it in the background — speaker-labeled, timestamped, and queryable like any other document. Supported formats: M4A, MP3, WAV, WebM, FLAC, OGG. Up to 500MB in Assistant, 4GB in Vault, two hours per file.You can also record directly from the Harvey mobile app for depositions or client meetings for up to two hours, and the transcript saves straight to Vault with each speaker automatically labeled. Only the transcript is retained, not the original audio.For law firms: Client calls, witness interviews, and internal case discussions can go straight from recording to searchable transcript, with no separate transcription step and no manual cleanup before analysis can start.For in-house legal teams: Recorded regulatory interviews, internal investigation interviews, or board and compliance calls become queryable Harvey documents in minutes rather than sitting in an audio file that has to be reviewed in real time.Video File TranscriptionThe same transcription works for video: MP4, MOV, AVI, DAV, WebM, up to two hours per file. For litigation teams, this means bulk-uploading a set of deposition videos into Vault and querying across all of them at once: surfacing contradictions, tracing a timeline, pulling every mention of an issue into a Review Table.For law firms: This is built for the reality of litigation, where video evidence is constant: depositions, witness interviews, surveillance footage. A team can bulk-upload dozens or even hundreds of deposition videos into Vault and query across the entire set at once surfacing contradictions between witnesses, tracing a timeline across statements, or pulling every reference to a specific issue into a Review Table.For in-house legal teams: Recorded interviews and video evidence sit alongside contracts and correspondence, and native video support means all of it can be reviewed in one platform instead of split across tools.Why This MattersWith image, audio, and video now natively understood, Harvey can support end-to-end analysis, research, and drafting across the full range of evidence a matter actually contains, not just the parts that happen to be text. That means fewer handoffs between tools, faster turnaround from raw file to usable answer, and a more complete view of the record. Legal work has always been multimodal. Now Harvey is too.Ready to see it in practice? Contact your account team or reach out to us below to get started.Request a DemoUnable to load form. Please try again.Try AgainThank you!We'll be in touch shortly.Next UpIntroducing Harvey IIThe Brief: August 2026A Smarter Inbox Built for Legal Work: The New Harvey for Outlook

归纳

音视频图像理解技术旨在使机器能够同时处理和分析音频、视频与图像等多模态数据,实现对场景、物体、动作及语义信息的综合感知。当前研究涵盖跨模态特征提取、对齐与融合、时序建模以及大模型驱动的感知推理等方向,广泛应用于内容审核、智能监控、自动驾驶及多媒体检索等领域。

点评

音视频图像理解技术的多模态数据聚合放大了训练数据来源合法性、生成内容标识义务及跨境部署中的算法问责风险。

AI 自动整理 · 未人工精选

法律视角点评

AI 生成 · 人工审核

核心关切

音视频图像理解技术的多模态数据聚合放大了训练数据来源合法性、生成内容标识义务及跨境部署中的算法问责风险。

法律依据

EU AI Act Article 50(2) (Regulation (EU) 2024/1689) 对生成或操控合成音视频图像内容的披露义务作出规定。

实务启示

中国法律人可借鉴该条款为出海多模态产品建立数据来源尽调、内容水印与元数据披露机制,并在协议和隐私政策中嵌入可解释性与问责条款。

查看原文
音视频图像理解技术研究进展