Chuyển đến nội dung chính

Bài đăng

Hiển thị các bài đăng có nhãn RAG

Chọn mô hình embedding: những câu hỏi thật sự quan trọng

Quy trình thường thấy là: mở bảng xếp hạng, sắp theo điểm trung bình, lấy mô hình đứng đầu vừa túi tiền, rồi đi tiếp. Nó cho ra lựa chọn bảo vệ được khoảng một nửa số lần, và ở nửa còn lại, nó hỏng một cách đắt đỏ — vì đổi mô hình embedding nghĩa là nhúng lại toàn bộ và dựng lại mọi chỉ mục. Đây là những câu hỏi dự báo kết quả tốt hơn thứ hạng. Nó có chạy tốt trên văn bản của bạn không? Điểm trung bình trên benchmark là một hỗn hợp có trọng số của nhiều tác vụ, mà phần lớn không phải tác vụ của bạn. Một mô hình dẫn đầu tổng thể có thể tụt hậu tệ hại ở đúng cái bạn cần — điều khoản pháp lý, mã hàng, ticket hỗ trợ khách hàng tiếng Việt, hay mã nguồn. Bài đánh giá thật sự quan trọng chỉ mất một buổi chiều: Thu 50-100 truy vấn thật từ log của bạn (hoặc tự viết, nếu chưa có lưu lượng). Với mỗi truy vấn, đánh dấu những tài liệu trong kho đáng lẽ phải được truy hồi. Việc gán nhãn này chính là phần công việc thật. Nhúng kho tài liệu bằng từng mô hình ứng viên, chạy các truy vấn, đo r...

Chia nhỏ tài liệu cho truy hồi: quyết định âm thầm chặn trần chất lượng RAG

Một hệ thống truy hồi trả lời đúng câu “thời hạn hoàn tiền của chúng ta là bao lâu?” nhưng thất bại với câu “thời hạn hoàn tiền có khác với khách doanh nghiệp không?” thường không có vấn đề về mô hình. Nó có một đoạn chứa chính sách chung và một đoạn khác, không bao giờ được truy hồi, chứa ngoại lệ. Việc chia đoạn quyết định một đơn vị truy hồi là gì . Làm sai và không lượng rerank, kỹ thuật prompt hay nâng cấp mô hình nào cứu lại được thông tin bạn đã cắt rời. Chia theo kích thước cố định là mốc so sánh, không phải chiến lược Cách mặc định ai cũng bắt đầu — cắt mỗi N ký tự với M ký tự chồng lấn — đáng được hiểu chính xác, vì đó là thứ bạn sẽ đem ra so sánh. def fixed_chunks (text, size = 1000 , overlap = 200 ): chunks, start = [], 0 while start < len (text): chunks.append(text[start:start + size]) start += size - overlap return chunks Nó làm đúng chỗ nào: đoạn đều nhau, chi phí đoán được, không giả định gì về định dạng tài liệu. Nó là...

Choosing an embedding model: the questions that actually matter

The usual process is: open a leaderboard, sort by average score, take the top model that fits the budget, move on. It produces a defensible choice roughly half the time, and the half where it fails, it fails expensively — because switching embedding models means re-embedding everything and rebuilding every index. Here are the questions that predict the outcome better than rank does. Does it work on your text? A benchmark average is a weighted mixture of tasks, most of which are not yours. A model that leads overall can trail badly on the one thing you need — legal clauses, product SKUs, Vietnamese customer support tickets, code. The evaluation that matters takes an afternoon: Collect 50-100 real queries from your logs (or write them, if you have no traffic yet). For each, mark which documents in your corpus should be retrieved. This labelling is the actual work. Embed the corpus with each candidate model, run the queries, measure recall@k for the k you actually feed to the mo...

Chunking for retrieval: the decision that quietly caps your RAG quality

A retrieval system that answers “what is our refund window?” correctly and fails on “does the refund window differ for enterprise customers?” usually does not have a model problem. It has a chunk that contains the general policy and a different chunk, retrieved never, that contains the exception. Chunking decides what a single retrievable unit is . Get it wrong and no amount of reranking, prompt engineering or model upgrading recovers the information you split apart. Fixed-size splitting is a baseline, not a strategy The default everyone starts with — split every N characters with M overlap — is worth understanding precisely, because it is the thing you will be comparing against. def fixed_chunks (text, size = 1000 , overlap = 200 ): chunks, start = [], 0 while start < len (text): chunks.append(text[start:start + size]) start += size - overlap return chunks What it gets right: uniform chunks, predictable cost, no assumptions about docum...

Choosing an embedding model: the questions that actually matter

The usual process is: open a leaderboard, sort by average score, take the top model that fits the budget, move on. It produces a defensible choice roughly half the time, and the half where it fails, it fails expensively — because switching embedding models means re-embedding everything and rebuilding every index. Here are the questions that predict the outcome better than rank does. Does it work on your text? A benchmark average is a weighted mixture of tasks, most of which are not yours. A model that leads overall can trail badly on the one thing you need — legal clauses, product SKUs, Vietnamese customer support tickets, code. The evaluation that matters takes an afternoon: Collect 50-100 real queries from your logs (or write them, if you have no traffic yet). For each, mark which documents in your corpus should be retrieved. This labelling is the actual work. Embed the corpus with each candidate model, run the queries, measure recall@k for the k you actually feed to the mo...

Chunking for retrieval: the decision that quietly caps your RAG quality

A retrieval system that answers “what is our refund window?” correctly and fails on “does the refund window differ for enterprise customers?” usually does not have a model problem. It has a chunk that contains the general policy and a different chunk, retrieved never, that contains the exception. Chunking decides what a single retrievable unit is . Get it wrong and no amount of reranking, prompt engineering or model upgrading recovers the information you split apart. Fixed-size splitting is a baseline, not a strategy The default everyone starts with — split every N characters with M overlap — is worth understanding precisely, because it is the thing you will be comparing against. def fixed_chunks (text, size = 1000 , overlap = 200 ): chunks, start = [], 0 while start < len (text): chunks.append(text[start:start + size]) start += size - overlap return chunks What it gets right: uniform chunks, predictable cost, no assumptions about docum...

Retrieval không nói dối: dựng eval trước khi dựng RAG

Bug report lúc nào cũng cùng một dạng. “Trợ lý bịa ra thời hạn hoàn tiền của mình.” Ai đó mở trace, thấy một đoạn văn sai nhưng đầy tự tin, rồi quy lỗi cho model. Prompt được thêm một đoạn hướng dẫn nữa. Có khi temperature bị hạ xuống. Có khi ai đó đề xuất đổi model. Rồi bạn nhìn lên một tầng và thấy retriever đã đưa cho model ba chunk: chính sách vận chuyển, phần mở đầu của điều khoản dịch vụ, và một chunk đứt giữa câu ngay trước chỗ nêu thời hạn hoàn tiền. Model không hẳn là hallucinate — nó ứng biến trên một lỗ hổng. Không prompt nào sửa được chuyện đó. Đổi model cũng không: một model tốt hơn với đúng ba chunk ấy sẽ cho ra một đoạn sai thuyết phục hơn . Đây là dạng lỗi RAG phổ biến nhất, và bạn không thể nhìn ra nó từ output. Generation là phần duy nhất bạn đọc, nên generation là phần bạn đổ lỗi. Cách sửa là một phép đo tách đôi hai nửa, dựng trước khi bạn tinh chỉnh bất cứ thứ gì — nếu không, mọi thay đổi bạn làm đều là tung đồng xu mà không chấm điểm được. Dưới đây là phiên bản...