基于湖库一体架构,统一管理结构化、半结构化与非结构化等多模态数据,一个系统承载事务处理、实时分析与 AI 工作负载。
OB Cloud Vector 与 LangChain 集成
更新时间:2026-05-14 14:28:04
OceanBase 从 V4.3.3 开始支持向量类型存储、向量索引、embedding 向量检索的能力。可以将向量化后的数据存储在 OceanBase,供下一步的检索使用。
LangChain 是一个用于开发由语言模型驱动的应用程序的框架。它使得应用程序能够:
具有上下文感知能力:将语言模型连接到上下文来源(提示指令,少量的示例,需要回应的内容等)。
具有推理能力:依赖语言模型进行推理(根据提供的上下文如何回答,采取什么行动等)。
本文结合通义千问 API,介绍如何将 OceanBase Cloud 中的向量检索功能、通义千问与 LangChain 集成实现文档问答。
前提条件
- 您的环境中有可用的事务型(MySQL)集群实例。
请参考 创建租户完成租户创建后,再参考下述步骤操作。
您的环境中已存在可以使用的 MySQL 兼容模式租户和 MySQL 数据库和账号,并已对数据库账号授予读写权限。如需创建,详情参见 创建账号和 创建数据库(仅 MySQL)。
您拥有项目管理员或实例管理员角色可对项目中的实例进行读写操作,如无权限,可联系组织管理员添加权限。
安装 Python 3.9 及以上版本 和相应 pip。如果您的机器上 Python 版本较低,可以使用 Miniconda 来创建新的 Python 3.9 及以上的环境,具体可参考 Miniconda 安装指南。
安装依赖。
python3 -m pip install -U langchain-oceanbase python3 -m pip install langchain_community python3 -m pip install dashscope
步骤一:获取数据库连接信息
在下拉框中,根据 ID 选择您的集群实例。
进入 实例工作台 页面。
单击连接,选择 获取连接串。
在弹出框中选择 使用公共网络。
获取访问地址,选择 添加当前浏览器 IP 地址。
填写数据库相关信息,复制连接串。
步骤二:注册 LLM 平台账号
注册阿里云百炼账号,开通模型服务并获取 API 密钥。
注意
开通阿里云百炼大模型服务需要您跳转至第三方平台完成。此操作将遵循第三方平台的收费规则,并可能产生相应费用。请在继续前,访问其官网或查阅相关文档,确认并接受其收费标准。如不同意,请勿继续操作。




配置 API Key 到环境变量:
export DASHSCOPE_API_KEY="YOUR_DASHSCOPE_API_KEY"
步骤三:构建您的 AI 助手
加载并分割文档
下载示例数据,将其拆分成每块约 1000 个字符的块 CharacterTextSplitter。
from langchain_community.document_loaders import TextLoader
from langchain_community.embeddings import DashScopeEmbeddings
from langchain_text_splitters import CharacterTextSplitter
from langchain_oceanbase.vectorstores import OceanbaseVectorStore
import os
import requests
DASHSCOPE_API = os.environ.get("DASHSCOPE_API_KEY", "")
embeddings = DashScopeEmbeddings(
model="text-embedding-v1", dashscope_api_key=DASHSCOPE_API
)
url = "https://raw.githubusercontent.com/GITHUBear/langchain/refs/heads/master/docs/docs/how_to/state_of_the_union.txt"
res = requests.get(url)
with open("state_of_the_union.txt", "w") as f:
f.write(res.text)
loader = TextLoader('./state_of_the_union.txt')
documents = loader.load()
text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=0)
docs = text_splitter.split_documents(documents)
将数据插入 OB Cloud 中
connection_args = {
"host": "127.0.0.1",
"port": "2881",
"user": "root@sun",
"password": "",
"db_name": "test",
}
DEMO_TABLE_NAME = "demo_ann"
ob = OceanbaseVectorStore(
embedding_function=embeddings,
table_name=DEMO_TABLE_NAME,
connection_args=connection_args,
drop_old=True,
normalize=True,
)
res = ob.add_documents(documents=docs)
向量搜索
此步骤演示如何从文档 state_of_the_union.txt 中查询 “What did the president say about Ketanji Brown Jackson”。
query = "What did the president say about Ketanji Brown Jackson"
docs_with_score = ob.similarity_search_with_score(query, k=3)
for doc, score in docs_with_score:
print("-" * 80)
print("Score: ", score)
print(doc.page_content)
print("-" * 80)
预期输出如下:
--------------------------------------------------------------------------------
Score: 1.204783671324283
Tonight. I call on the Senate to: Pass the Freedom to Vote Act. Pass the John Lewis Voting Rights Act. And while you’re at it, pass the Disclose Act so Americans can know who is funding our elections.
Tonight, I’d like to honor someone who has dedicated his life to serve this country: Justice Stephen Breyer—an Army veteran, Constitutional scholar, and retiring Justice of the United States Supreme Court. Justice Breyer, thank you for your service.
One of the most serious constitutional responsibilities a President has is nominating someone to serve on the United States Supreme Court.
And I did that 4 days ago, when I nominated Circuit Court of Appeals Judge Ketanji Brown Jackson. One of our nation’s top legal minds, who will continue Justice Breyer’s legacy of excellence.
--------------------------------------------------------------------------------
--------------------------------------------------------------------------------
Score: 1.2146663629717394
It is going to transform America and put us on a path to win the economic competition of the 21st Century that we face with the rest of the world—particularly with China.
As I’ve told Xi Jinping, it is never a good bet to bet against the American people.
We’ll create good jobs for millions of Americans, modernizing roads, airports, ports, and waterways all across America.
And we’ll do it all to withstand the devastating effects of the climate crisis and promote environmental justice.
We’ll build a national network of 500,000 electric vehicle charging stations, begin to replace poisonous lead pipes—so every child—and every American—has clean water to drink at home and at school, provide affordable high-speed internet for every American—urban, suburban, rural, and tribal communities.
4,000 projects have already been announced.
And tonight, I’m announcing that this year we will start fixing over 65,000 miles of highway and 1,500 bridges in disrepair.
--------------------------------------------------------------------------------
--------------------------------------------------------------------------------
Score: 1.2193955178945004
Vice President Harris and I ran for office with a new economic vision for America.
Invest in America. Educate Americans. Grow the workforce. Build the economy from the bottom up
and the middle out, not from the top down.
Because we know that when the middle class grows, the poor have a ladder up and the wealthy do very well.
America used to have the best roads, bridges, and airports on Earth.
Now our infrastructure is ranked 13th in the world.
We won’t be able to compete for the jobs of the 21st Century if we don’t fix that.
That’s why it was so important to pass the Bipartisan Infrastructure Law—the most sweeping investment to rebuild America in history.
This was a bipartisan effort, and I want to thank the members of both parties who worked to make it happen.
We’re done talking about infrastructure weeks.
We’re going to have an infrastructure decade.
--------------------------------------------------------------------------------