開發者工具 / 31
RAG文字分塊與重疊視覺化工具
在瀏覽器中視覺化RAG向量索引的文字分割策略、token/字元分塊大小與重疊內容。
RAG文字分塊與重疊視覺化工具: TOEA會依照您選擇的策略,在本機將文件分割成文字區塊。工具會計算字元範圍、token估算值(約4個字元/token),以及連續文字區塊之間明確重疊的文字。 全程在瀏覽器中於本機執行,檔案不會上傳至伺服器。
- 分類
- 開發者工具
- 使用次數
- 在瀏覽器中
- 費用
- 免費・免註冊
- 可用狀態
- 可立即使用
產生的文件分塊與重疊部分
# Introduction to Retrieval-Augmented Generation (RAG) Retrieval-Augmented Generation enhances Large Language Models by grounding outputs on external knowledge bases. Instead of relying solely on parametric memory learned during training, RAG models fetch relevant text passages at inference time.
Retrieval-Augmented Generation enhances Large Language Models by grounding outputs on external knowledge bases. Instead of relying solely on parametric memory learned during training, RAG models fetch relevant text passages at inference time. ## Chunking Strategies
## Chunking Strategies Proper text chunking is critical for RAG accuracy. Fixed-size chunking splits documents into uniform character intervals. Sentence-based chunking preserves grammatical units. Markdown chunking splits along section headers. ## Vector Indexing
Proper text chunking is critical for RAG accuracy. Fixed-size chunking splits documents into uniform character intervals. Sentence-based chunking preserves grammatical units. Markdown chunking splits along section headers. ## Vector Indexing
## Vector Indexing Once chunks are generated, embedding models convert text chunks into dense vector representations. These vectors are indexed in vector databases like Pinecone, Qdrant, or PGVector for fast cosine similarity search.
字元、Token與模型限制
這裡的區塊大小以字元計算,Token數量則假設每個Token約有四個字元,這對英文散文通常成立。程式碼、數字和許多其他語言每個字元會使用更多Token,因此請預留餘裕。請根據嵌入模型的輸入限制檢查結果:許多BERT類型模型會在512個Token處停止,通常會捨棄其餘內容;OpenAI的text-embedding-3模型則接受8,191個Token。在2,000字元的上限下,一個區塊約有500個英文Token。
閱讀預覽
請留意從想法中途開始的區塊、與其介紹文字分開的標題,以及被切成兩半的表格或程式碼區塊;這些內容在擷取時都不會帶有賦予其意義的上下文。使用句子、段落和Markdown策略時,若單一單位長度超過區塊大小,仍會完整保留,因此過大的區塊通常表示其中有很長的段落或區段。
使用這三種策略時,重疊內容也會以完整的句子、段落或區段為單位回退,而不是依照精確的字元數。許多處理流程也會在每個區塊嵌入前,先加上文件標題或區段標題。
使用方式
- 將文件或程式碼貼到編輯器中。
- 選擇分割策略(段落、句子、Markdown標題或固定大小),並調整目標分塊大小與重疊滑桿。
- 查看以顏色標示的分塊邊界、字元/token指標,以及重疊範圍。
隱私與限制
您的來源文件會100%留在瀏覽器記憶體中。
相關工具
常見問題
RAG應該使用哪種文字分塊策略?
一般文字最適合使用段落邊界;結構化文件適合使用Markdown標題;需要高精準度問答時,則適合使用句子邊界。
為什麼文字區塊重疊很重要?
重疊可避免文字區塊邊界造成語意遺失,確保向量相似度搜尋能擷取跨越分割點的上下文。
免費工具・使用在瀏覽器中次・免註冊