代码分割

CodeTextSplitter允许您使用多种语言支持来分割代码。导入枚举Language并指定语言。

from langchain.text_splitter import (
    RecursiveCharacterTextSplitter,
    Language,
)

支持的语言列表

[e.value for e in Language]

['cpp',
 'go',
 'java',
 'js',
 'php',
 'proto',
 'python',
 'rst',
 'ruby',
 'rust',
 'scala',
 'swift',
 'markdown',
 'latex',
 'html',
 'sol']

您还可以查看给定语言使用的分隔符

RecursiveCharacterTextSplitter.get_separators_for_language(Language.PYTHON)

['\nclass ', '\ndef ', '\n\tdef ', '\n\n', '\n', ' ', '']

Python

以下是使用PythonTextSplitter的示例

PYTHON_CODE = """
def hello_world():
    print("Hello, World!")

# 调用函数
hello_world()
"""
python_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.PYTHON, chunk_size=50, chunk_overlap=0
)
python_docs = python_splitter.create_documents([PYTHON_CODE])
python_docs

[Document(page_content='def hello_world():\n    print("Hello, World!")', metadata={}),
 Document(page_content='# 调用函数\nhello_world()', metadata={})]

JS

以下是使用JS文本分割器的示例

JS_CODE = """
function helloWorld() {
  console.log("Hello, World!");
}

// 调用函数
helloWorld();
"""

js_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.JS, chunk_size=60, chunk_overlap=0
)
js_docs = js_splitter.create_documents([JS_CODE])
js_docs

[Document(page_content='function helloWorld() {\n  console.log("Hello, World!");\n}', metadata={}),
 Document(page_content='// 调用函数\nhelloWorld();', metadata={})]

Markdown

以下是使用Markdown文本分割器的示例。

markdown_text = """
# 🦜️🔗 LangChain

⚡ Building applications with LLMs through composability ⚡

## Quick Install

```bash
# Hopefully this code block isn't split
pip install langchain


As an open source project in a rapidly developing field, we are extremely open to contributions.

md_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.MARKDOWN, chunk_size=60, chunk_overlap=0
)
md_docs = md_splitter.create_documents([markdown_text])
md_docs

[Document(page_content='# 🦜️🔗 LangChain', metadata={}),
 Document(page_content='⚡ 通过组合性构建LLM应用程序 ⚡', metadata={}),
 Document(page_content='## 快速安装', metadata={}),
 Document(page_content="```bash\n# 希望这个代码块不会被分割", metadata={}),
 Document(page_content='pip install langchain', metadata={}),
 Document(page_content='```', metadata={}),
 Document(page_content='作为一个快速发展的领域中的开源项目，我们', metadata={}),
 Document(page_content='非常欢迎贡献。', metadata={})]

Latex

以下是关于Latex文本的示例

latex_text = """
\documentclass{article}

\begin{document}

\maketitle

\section{Introduction}
Large language models (LLMs) are a type of machine learning model that can be trained on vast amounts of text data to generate human-like language. In recent years, LLMs have made significant advances in a variety of natural language processing tasks, including language translation, text generation, and sentiment analysis.

\subsection{History of LLMs}
The earliest LLMs were developed in the 1980s and 1990s, but they were limited by the amount of data that could be processed and the computational power available at the time. In the past decade, however, advances in hardware and software have made it possible to train LLMs on massive datasets, leading to significant improvements in performance.

\subsection{Applications of LLMs}
LLMs have many applications in industry, including chatbots, content creation, and virtual assistants. They can also be used in academia for research in linguistics, psychology, and computational linguistics.

\end{document}
"""

latex_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.MARKDOWN, chunk_size=60, chunk_overlap=0
)
latex_docs = latex_splitter.create_documents([latex_text])
latex_docs

[Document(page_content='\\documentclass{article}\n\n\\begin{document}\n\n\\maketitle', metadata={}),
 Document(page_content='\\section{Introduction}', metadata={}),
 Document(page_content='Large language models (LLMs) are a type of machine learning', metadata={}),
 Document(page_content='model that can be trained on vast amounts of text data to', metadata={}),
 Document(page_content='generate human-like language. In recent years, LLMs have', metadata={}),
 Document(page_content='made significant advances in a variety of natural language', metadata={}),
 Document(page_content='processing tasks, including language translation, text', metadata={}),
 Document(page_content='generation, and sentiment analysis.', metadata={}),
 Document(page_content='\\subsection{History of LLMs}', metadata={}),
 Document(page_content='The earliest LLMs were developed in the 1980s and 1990s,', metadata={}),
 Document(page_content='but they were limited by the amount of data that could be', metadata={}),
 Document(page_content='processed and the computational power available at the', metadata={}),
 Document(page_content='time. In the past decade, however, advances in hardware and', metadata={}),
 Document(page_content='software have made it possible to train LLMs on massive', metadata={}),
 Document(page_content='datasets, leading to significant improvements in', metadata={}),
 Document(page_content='performance.', metadata={}),
 Document(page_content='\\subsection{Applications of LLMs}', metadata={}),
 Document(page_content='LLMs have many applications in industry, including', metadata={}),
 Document(page_content='chatbots, content creation, and virtual assistants. They', metadata={}),
 Document(page_content='can also be used in academia for research in linguistics,', metadata={}),
 Document(page_content='psychology, and computational linguistics.', metadata={}),
 Document(page_content='\\end{document}', metadata={})]

HTML

以下是使用HTML文本分割器的示例

html_text = """
<!DOCTYPE html>
<html>
    <head>
        <title>🦜️🔗 LangChain</title>
        <style>
            body {
                font-family: Arial, sans-serif;
            }
            h1 {
                color: darkblue;
            }
        </style>
    </head>
    <body>
        <div>
            <h1>🦜️🔗 LangChain</h1>
            <p>⚡ Building applications with LLMs through composability ⚡</p>
        </div>
        <div>
            As an open source project in a rapidly developing field, we are extremely open to contributions.
        </div>
    </body>
</html>

html_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.HTML, chunk_size=60, chunk_overlap=0
)
html_docs = html_splitter.create_documents([html_text])
html_docs

[Document(page_content='<!DOCTYPE html>\n<html>', metadata={}),
 Document(page_content='<head>\n        <title>🦜️🔗 LangChain</title>', metadata={}),
 Document(page_content='<style>\n            body {\n                font-family: Aria', metadata={}),
 Document(page_content='l, sans-serif;\n            }\n            h1 {', metadata={}),
 Document(page_content='color: darkblue;\n            }\n        </style>\n    </head', metadata={}),
 Document(page_content='>', metadata={}),
 Document(page_content='<body>', metadata={}),
 Document(page_content='<div>\n            <h1>🦜️🔗 LangChain</h1>', metadata={}),
 Document(page_content='<p>⚡ Building applications with LLMs through composability ⚡', metadata={}),
 Document(page_content='</p>\n        </div>', metadata={}),
 Document(page_content='<div>\n            As an open source project in a rapidly dev', metadata={}),
 Document(page_content='eloping field, we are extremely open to contributions.', metadata={}),
 Document(page_content='</div>\n    </body>\n</html>', metadata={})]

Solidity

以下是使用Solidity文本分割器的示例

SOL_CODE = """
pragma solidity ^0.8.20;
contract HelloWorld {
   function add(uint a, uint b) pure public returns(uint) {
       return a + b;
   }
}
"""

sol_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.SOL, chunk_size=128, chunk_overlap=0
)
sol_docs = sol_splitter.create_documents([SOL_CODE])
sol_docs

[
    Document(page_content='pragma solidity ^0.8.20;', metadata={}),
    Document(page_content='contract HelloWorld {\n   function add(uint a, uint b) pure public returns(uint) {\n       return a + b;\n   }\n}', metadata={})
]

代码分割

支持的语言列表

您还可以查看给定语言使用的分隔符

Python​

JS​

Markdown​

Latex​

HTML​

Solidity​

Python

JS

Markdown

Latex

HTML

Solidity