Chuyển đến nội dung chính

Bài đăng

Hiển thị các bài đăng có nhãn Production

Vision models in production: the parts the demo skips

The demo works on the first try: photograph a receipt, get back structured JSON with the merchant, date and total. Then you ship it, and users send you a receipt photographed at an angle in bad light with a thumb over the total, a 40-page scanned contract, a screenshot of a spreadsheet, and a blurry photo of another screen showing a receipt. The model handles more of that than you would expect, and the parts it does not handle are the ones that determine whether the feature is usable. Resolution is the cost dial Images become tokens, and the count scales with pixel area. The precise formula differs by provider, but the shape is universal: a large image can cost more than a page of text, and most providers downscale images above a maximum dimension before processing anyway. Two consequences follow. Sending a full-resolution phone photo is usually waste. An 12-megapixel image gets downscaled by the provider, so you paid to upload pixels that were discarded. Resize before sending: ...

Flutter error handling: catching what actually reaches users

A crash-free rate of 99.8% sounds excellent until you realise it only counts crashes your reporting tool was wired to see. In Flutter, an error can escape through four different doors, and most apps only guard one or two of them. The four doors Door Catches Missed if unwired FlutterError.onError Errors inside the framework: build, layout, paint, gesture callbacks Red screens, silent layout failures PlatformDispatcher.instance.onError Uncaught async errors in the root zone Most Future failures Isolate.current.addErrorListener Errors in isolates you spawned Every background-compute failure Native crash handler Platform-level crashes, plugin native code Anything that kills the process Here is all four, wired once at startup: Future < void > main () async { WidgetsFlutterBinding . ensureInitialized (); await Firebase . initializeApp (); final crashlytics = FirebaseCrashlytics .instance; // 1. Framework errors. FlutterError .onEr...

Vision models in production: the parts the demo skips

The demo works on the first try: photograph a receipt, get back structured JSON with the merchant, date and total. Then you ship it, and users send you a receipt photographed at an angle in bad light with a thumb over the total, a 40-page scanned contract, a screenshot of a spreadsheet, and a blurry photo of another screen showing a receipt. The model handles more of that than you would expect, and the parts it does not handle are the ones that determine whether the feature is usable. Resolution is the cost dial Images become tokens, and the count scales with pixel area. The precise formula differs by provider, but the shape is universal: a large image can cost more than a page of text, and most providers downscale images above a maximum dimension before processing anyway. Two consequences follow. Sending a full-resolution phone photo is usually waste. An 12-megapixel image gets downscaled by the provider, so you paid to upload pixels that were discarded. Resize before sending: ...

Đưa một tính năng LLM lên production: bản demo chỉ là 20% công việc

Bản demo lúc nào cũng chạy ngon. Bạn dán một input đẹp vào, model trả về thứ gì đó ấn tượng, bạn đem khoe ở standup, cả nhóm đồng ý là phải ship. Hai tuần sau nó nằm trên production và bạn đang đọc một ticket hỗ trợ trong đó model nói với khách rằng yêu cầu hoàn tiền của họ đã được duyệt. Bản demo là 20% dễ. 80% còn lại là tất cả những thứ ngăn một hàm xác suất kéo sập ứng dụng của bạn — không phải vì model tệ, mà vì bạn vừa cắm một thành phần không tất định vào một hệ thống được thiết kế trên giả định rằng hàm trả về đúng cái mà chữ ký của nó khai báo. Dưới đây là thứ tự tôi thật sự làm. Thứ tự quan trọng hơn từng kỹ thuật riêng lẻ, vì mỗi bước cho bạn biết bước kế tiếp có đáng làm hay không. Bỏ qua bộ eval thì bạn sẽ mất ba ngày tinh chỉnh prompt mà không có cách nào biết mình có cải thiện được gì không. Bỏ qua guardrail thì bạn sẽ biết về lỗi format output của mình qua lời người dùng. Không có gì ở đây gắn với một nhà cung cấp hay một framework cụ thể. Đây vẫn là kỷ luật bạn áp c...

Shipping an LLM feature: the demo is 20% of the work

The demo always works. You paste a good input, the model returns something impressive, you show it in standup, and everyone agrees it should ship. Two weeks later it is in production and you are reading a support ticket where the model told a customer their refund was approved. The demo is the easy 20%. The other 80% is everything that stops a probabilistic function from taking your application down with it — not because the model is bad, but because you wired a non-deterministic component into a system that was designed on the assumption that functions return what their signature says. What follows is the order I actually build these things in. The order matters more than any individual technique, because each step tells you whether the next one is even worth doing. If you skip the eval set you will spend three days tuning a prompt and have no way to know if you improved it. If you skip the guardrails you will find out about your output-format bug from a user. None of this is speci...