VOL. 06 / 编辑型法律 AI 情报本周信号已校对 · 06
← 返回首页
ARTICLE DOSSIERREF / C123F92B
媒体🌐 国际重要

Gemma 2、Llama与Swallow模型法律现状

SOURCE / DeepMind-官方网站 · gemma 2 llama swallow

原文

gemma 2 llama swallow

完整原文

Institute of Science Tokyo creates powerful Japanese-focused LLM with Gemma 2Institute of Science Tokyo together with National Institute of Advanced Industrial Science and Technology (AIST) is working to create large language models (LLMs) that excel in Japanese. Their research consists of finding and deconstructing popular LLMs that demonstrate excellent capabilities in language understanding, generation, and dialogue, and pre-training them with large bodies of Japanese text to create their Swallow line of LLMs.Latest research efforts resulted in the creation of Gemma-2-Llama Swallow, a new LLM that delivers unparalleled Japanese language knowledge and performance, thanks also to Gemma’s high base-level proficiency in the language.The challengeThe institute recognized that many of the world’s most-popular LLMs focus on western languages like English, and lack reliable utility in European languages, southeastern languages, and in this case, Japanese. And the cost of the models that did have Japanese functionality outweighed the performance they offered.The Swallow developer team began creating Japanese-focused iterations of popular models from Llama, Mistral, and Mixtral to varying degrees of success. Eventually, the team chose Gemma 2 for its stronger base-level Japanese capabilities. “Gemma 2 already exhibited strong instruction-following and dialogue capabilities in Japanese,” said Naoaki Okazaki, professor at Institute of Science Tokyo. But the team knew they could make the model even better.Chart representing Gemma-2-Llama Swallow 27B IT v0.1 superior performance.The solutionTo improve Gemma 2’s Japanese proficiency, the team had to pre-train the model with a massive amount of Japanese training data specifically developed by the Swallow team. Gemma 2’s base-level proficiency in Japanese helped simplify the tokenizing process, allowing the team to skip modifying the tokenizer or token embeddings.Compared to other overseas LLMs, Gemma 2's tokenizer vocabulary includes a larger number of Japanese characters and words. This eliminated the need to modify the tokenizer or token embeddings before continual pre-training. Furthermore, its less restrictive licensing allowed us to leverage it for tasks such as filtering Japanese pre-training data and synthesizing instruction-tuning data. Professor Naoaki Okazaki Institute of Science Tokyo This team’s efforts resulted in the creation of Gemma-2-Llama Swallow in 2B, 9B, and 27B parameter versions. “Since Gemma 2 already exhibited strong instruction-following and dialogue capabilities in Japanese, we were able to employ imitation learning for the instruction tuning of our model from Gemma 2 27B,” said Professor Naoaki Okazaki. The size and performance of Gemma 2 27B helped the team save valuable resources as well.The impactTo put Gemma-2-Llama Swallow to the test, the team used 10 Japanese understanding and generation tasks, 10 English understanding and generation tasks, and the Japanese MT-bench to evaluate the performance of the models.Because Gemma 2 27B demonstrates performance comparable to other open LLMs in the 70B class, we were able to construct synthetic data using fewer computational resources. Professor Naoaki Okazaki Institute of Science Tokyo The team found that - at the time of release - Gemma-2-Llama Swallow demonstrated the highest performance among LLMs of comparable size in Japanese language understanding and generation tasks, while Gemma-2-Llama Swallow’s 2B and 9B variants stood out the most by exhibiting performance on par with LLMs one size class larger for less resources.What’s nextInstitute of Science Tokyo will continue to refine Gemma-2-Llama Swallow following its initial launch in May 2025. The team expects the LLM will encourage more research and adoption of Japanese proficient models. “Despite being relatively smaller LLMs, they could be applied to a variety of applications,” said Okazaki, highlighting the nimbleness of the models while simultaneously matching the performance of 70B-class LLMs. The team is also working on improving the Japanese capability of Gemma 3 to create Swallow models that are even faster, more powerful, and more cost-efficient.These models represent another step towards human-like intelligence in computers for the researchers at Institute of Science Tokyo. “Realizing artificial intelligence has been a dream since the dawn of computing,” said Okazaki. “Large language models are bringing us closer to that reality. We are entering an exciting era where, as AI developers, we can witness computers becoming increasingly intelligent.”More from the Gemmaverse View all The Ministry of Economy, Ecology and Agriculture of Ukraine digitizes licensing process with Gemma Learn more Adaptive ML trains Gemma 3 for exceptional multilingual results Learn more Quarks improves user experiences with Gemma 2 and Gemma 3 Learn more Sarvam AI built a translation model with Gemma 3 to translate all 22 officially recognized Indian languages Learn more Institute of Science Tokyo creates powerful Japanese-focused LLM with Gemma 2 Learn more Introducing GAIA, a Brazilian Portuguese Gemma 3 model developed with ABRIA, CEIA, Nama, and Amadeus AI Learn more

归纳

Gemma 2为谷歌推出的开源大语言模型,Llama由Meta发布,而Swallow是专为日语设计的大模型。这些模型均面临数据合规与版权风险,开发者需遵循相应开源许可证条款。法律专家呼吁建立更明确的训练数据合法性框架及责任分配机制。目前,多个司法管辖区正研究AI生成内容的版权问题,各模型开发者已调整使用条款以应对欧盟AI法案等新规。社区也在讨论开源模式下训练数据侵权责任边界。

点评

开源AI模型在训练数据版权侵权与开源许可证合规之间面临双重不确定,当前责任分配机制尚属空白。

AI 自动整理 · 未人工精选

法律视角点评

AI 生成 · 人工审核

核心关切

开源AI模型在训练数据版权侵权与开源许可证合规之间面临双重不确定,当前责任分配机制尚属空白。

实务启示

中国法律人应结合《生成式人工智能服务管理暂行办法》,重点审查境外开源模型训练数据的合法性与许可证的传染性条款。

查看原文
Gemma 2、Llama与Swallow模型法律现状