Chuyển đến nội dung chính

Bài đăng

Hiển thị các bài đăng có nhãn Developer Tools

AI code review không ai muốn tắt: vấn đề là độ chính xác, không phải năng lực

Kịch bản hỏng lúc nào cũng giống nhau. Thứ Hai có người cắm một con AI reviewer vào CI. Nó để lại ba mươi comment trên pull request đầu tiên: hai mươi bảy cái là đặt tên biến, “cân nhắc tách đoạn này ra helper”, và một gợi ý thêm test mà cái test đó đã nằm sẵn ở file cách đó hai bậc. Ba cái là thật, và một trong ba là bug mất dữ liệu thật sự nằm trong nhánh retry. Không ai thấy ba cái đó. Đến thứ Tư mọi người lướt qua phần review của bot để xuống phần review của người. Đến thứ Hai tuần sau nó thành một check không chặn merge với thông báo đã tắt — một tích hợp chết nhưng vẫn đốt token mỗi lần push. Con bot không dở trong việc tìm bug. Nó tìm ra bug rồi. Nó dở ở chỗ không biết im lặng , mà trong code review thì đó chính là toàn bộ công việc. Ngân sách của một reviewer không phải là compute, mà là mức độ sẵn lòng đọc tiếp của đồng đội — và niềm tin đó cập nhật rất nhanh. Nếu trong mười comment đầu tiên một developer đọc chỉ có một cái hữu ích, họ đã học được một quy tắc — lướt qua thôi...

Running a model on your own machine: the four cases where it actually wins

There are two confident answers to “should I run a model locally,” and both are wrong. One says local models are now good enough that paying for an API is a waste. The other says anything you can fit on a laptop is a toy and you should stop wasting your evening. The useful answer is that “local versus API” isn’t a question about models at all. It’s a question about a specific workload, and it turns on four properties of that workload: whether the data is allowed to leave the machine, how many tokens a month you push through, whether there is a network, and how much of your latency budget a round trip eats. If none of those four is binding, the API is almost certainly the right call, and the honest reason is capability — on genuinely hard reasoning, a frontier model you rent still beats a quantised model you own. What follows is the concession first, then the four cases where local is not a compromise but the correct answer, then the mechanics: how to work out whether a model fits in ...

AI code review nobody mutes: the problem is precision, not capability

The failure mode is always the same. Someone wires an AI reviewer into CI on a Monday. It posts thirty comments on the first pull request: twenty-seven are variable naming, “consider extracting this into a helper”, and a suggestion to add a test that already exists two files over. Three are real, and one of those is a genuine data-loss bug in a retry path. Nobody finds the three. By Wednesday people scroll past the bot’s review to reach the human one. By the following Monday it is a non-blocking check with notifications muted — a dead integration that still burns tokens on every push. The bot was not bad at finding bugs. It found the bug. It was bad at not saying things , and in code review that is the entire job. A reviewer’s budget is not compute, it is your teammates’ willingness to keep reading, and that belief updates fast. If the first ten comments a developer reads contain one useful one, they have learned a rule — skim it, it is usually nothing — and that rule is applied to ...

Using AI without burning cash: free tiers, caching, and routing

Most people try to cut their AI bill by picking a cheaper model. That’s the smallest of the levers available, and usually the one that costs the most in quality. The bill is set by three things, in this order: What you pay for that’s actually free. Several providers give away real capacity with no card. What you pay for twice. The same 20,000-token system prompt, resent on every request, at full price. What you pay frontier rates for that doesn’t need them. Classification, extraction, renaming, formatting. Fix them in that order. The first is free money, the second is usually a 5–10x reduction on the same model, and only then does model choice matter. Here’s what each one is actually worth right now, with numbers from the providers’ own docs rather than from memory. Tier 0: what’s genuinely free today Free tiers are real, and they’re bigger than most people assume. But the numbers that matter aren’t the daily totals everyone quotes — they’re the per-minute ones, because th...