微信扫码
与创始人交个朋友
我要投稿
文本分割器的基本工作原理:
定制文本分割器的两个主要轴向:
主要参数和功能:
def transformer_doc():# 加载待分割长文本 with open('sys_boss.txt',encoding='UTF-8') as f:state_of_the_union = f.read()text_splitter = RecursiveCharacterTextSplitter(chunk_size = 100,chunk_overlap= 20,length_function = len,add_start_index = True,)docs = text_splitter.create_documents([state_of_the_union])print(docs[0])print(docs[1])metadatas = [{"document": 1}, {"document": 2}]documents = text_splitter.create_documents([state_of_the_union, state_of_the_union], metadatas=metadatas)print(documents[0])
def spit_code():print([e.value for e in Language])html_text = """<!DOCTYPE html><html><head><title>?️? LangChain</title><style>body {font-family: Arial, sans-serif;}h1 {color: darkblue;}</style></head><body><div><h1>?️? LangChain</h1><p>⚡ Building applications with LLMs through composability ⚡</p></div><div>As an open source project in a rapidly developing field, we are extremely open to contributions.</div></body></html>"""html_splitter = RecursiveCharacterTextSplitter.from_language(language=Language.HTML, chunk_size=60, chunk_overlap=0)html_docs = html_splitter.create_documents([html_text])print(html_docs)
53AI,企业落地应用大模型首选服务商
产品:大模型应用平台+智能体定制开发+落地咨询服务
承诺:先做场景POC验证,看到效果再签署服务协议。零风险落地应用大模型,已交付160+中大型企业
2024-09-18
再见了,LangChain!
2024-09-18
LLMFarm功能汇总:揭秘您可能错过的亮点功能!
2024-09-18
基于LangGraph构建LLM Agent
2024-09-18
Multi-Agent实战:构建复杂的数据处理与可视化系统
2024-09-16
LlamaIndex最新报告:构建高级大模型助手-RAG只是起点
2024-09-16
Tulip Agent:一种利用增删改查让LLM使用大量工具解决复杂任务的新框架!
2024-09-14
Multi Agent 多agent协同,不仅好玩还很实用,给你一个完整的demo
2024-09-13
【一文读懂】搭建RAG的万能工具LangChain
2024-04-11
2024-05-19
2024-04-12
2024-07-18
2024-04-08
2024-04-08
2024-03-31
2024-05-15
2024-06-03
2024-04-28
2024-08-27
2024-08-18
2024-08-16
2024-08-04
2024-07-31
2024-07-29
2024-07-28
2024-07-27