解析非结构化数据
Unstructured 处理非结构化数据
非结构化数据包括电子邮件、文档、图片、视频等没有预定义的数据模型或结构的数据类型
https://js.langchain.com/docs/how_to/document_loader_html
https://docs.unstructured.io/platform-api/legacy-api/free-api 获取api key
设置环境变量
export UNSTRUCTURED_API_KEY="..." export UNSTRUCTURED_API_URL="https://ip:8000/general/v0/general"
启动API服务
docker run -p 8000:8000 -d --rm --name unstructured-api downloads.unstructured.io/unstructured-io/unstructured-api:latest --port 8000 --host 0.0.0.0
解析示例
import { UnstructuredLoader } from "@langchain/community/document_loaders/fs/unstructured"; const filePath = "NCEPGDAS0P25.docx"; const loader = new UnstructuredLoader(filePath, { apiKey: '', apiUrl: 'http://XX.XX.XX.XX:8000/general/v0/general', }); const data = await loader.load(); console.log(data);