Chuyển đến nội dung chính

Bài đăng

Hiển thị các bài đăng có nhãn API

API versioning strategies for mobile clients — hướng dẫn Flutter

Old app binaries live for years. Design APIs for that reality. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Deprecation windows, feature negotiation, and forced upgrades. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Negotiate features via headers or capabilities endpoints. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure...

API versioning strategies for mobile clients

Old app binaries live for years. Design APIs for that reality. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Deprecation windows, feature negotiation, and forced upgrades. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Negotiate features via headers or capabilities endpoints. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates t...

API versioning strategies for mobile clients — hướng dẫn Flutter

Old app binaries live for years. Design APIs for that reality. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Deprecation windows, feature negotiation, and forced upgrades. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Negotiate features via headers or capabilities endpoints. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure...

API versioning strategies for mobile clients

Old app binaries live for years. Design APIs for that reality. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Deprecation windows, feature negotiation, and forced upgrades. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Negotiate features via headers or capabilities endpoints. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates t...

Prompt caching thực chiến: thắng thua nằm ở thứ tự prompt, không nằm ở cái flag

Lần đầu bật prompt caching, đa số mọi người thêm một field vào request, deploy, rồi thấy… không có gì thay đổi. Hoá đơn y hệt. Latency y hệt. Tính năng đã bật và nó không làm gì cả. Đó là kết quả bình thường, và không phải bug. Prompt caching không phải một cái công tắc kiểu “tái sử dụng được gì thì tái sử dụng”. Nó là so khớp prefix trên đúng từng byte của prompt sau khi render . Provider băm prompt của bạn từ byte đầu tiên trở đi và tìm một entry đã lưu để nối tiếp. Nếu byte thứ 47 trong prompt 30.000 token của bạn khác lần trước — vì đó là một cái đồng hồ — thì không còn prefix nào tái dùng được, và 29.900 token còn lại bị xử lý lại từ đầu với giá đầy đủ. Nên cái flag không phải là phần việc. Phần việc là sắp lại prompt sao cho phần bất biến nằm trước về mặt vật lý, còn phần theo từng request nằm sau cùng. Đó là thay đổi trong code dựng prompt, không phải trong lời gọi API. Khi thứ tự đã đúng, cái flag gần như là tiền cho không. Khi thứ tự còn sai, có rắc bao nhiêu cache marker cũ...

Prompt caching in practice: the win is in the ordering, not the flag

The first time most people enable prompt caching, they add one field to the request, deploy, and see nothing change. The bill is the same. Latency is the same. The feature is on and it does nothing. That’s the normal outcome, and it isn’t a bug. Prompt caching is not a per-request switch that says “reuse whatever you can.” It is a prefix match on the exact bytes of the rendered prompt . The provider hashes your prompt from the very first byte forward and looks for a stored entry it can resume from. If byte 47 of your 30,000-token prompt is different from last time — because it’s a clock reading — there is no reusable prefix, and the other 29,900 tokens get processed from scratch at full price. So the flag is not the work. The work is restructuring the prompt so the invariant part is physically first and the per-request part is physically last. That’s a change to your prompt-assembly code, not to your API call. Once the ordering is right, the flag is nearly free money. While the order...

Dùng AI mà không đốt tiền: free tier, caching và định tuyến model

Phần lớn mọi người cắt hóa đơn AI bằng cách đổi sang model rẻ hơn. Đó là đòn bẩy nhỏ nhất trong đám, và thường là đòn bẩy trả giá đắt nhất bằng chất lượng. Hóa đơn được quyết định bởi ba thứ, theo đúng thứ tự này: Thứ bạn đang trả tiền dù nó vốn miễn phí. Vài nhà cung cấp cho hạn mức thật, không cần thẻ. Thứ bạn trả tiền hai lần. Cùng một system prompt 20.000 token, gửi lại nguyên giá ở mỗi request. Thứ bạn trả giá frontier trong khi nó không cần. Phân loại, trích xuất, đổi tên, định dạng. Sửa theo đúng thứ tự đó. Cái đầu là tiền cho không, cái thứ hai thường giảm 5–10 lần trên cùng một model, và chỉ sau đó việc chọn model mới đáng bàn. Dưới đây là giá trị thật của từng cái ngay lúc này, với số liệu lấy từ tài liệu chính chủ chứ không phải từ trí nhớ. Tier 0: hôm nay cái gì thật sự miễn phí Free tier là thật, và rộng hơn phần lớn mọi người tưởng. Nhưng con số quan trọng không phải cái tổng theo ngày mà ai cũng đem đi khoe — mà là con số theo phút , vì đó mới là thứ chặn bạn...

Using AI without burning cash: free tiers, caching, and routing

Most people try to cut their AI bill by picking a cheaper model. That’s the smallest of the levers available, and usually the one that costs the most in quality. The bill is set by three things, in this order: What you pay for that’s actually free. Several providers give away real capacity with no card. What you pay for twice. The same 20,000-token system prompt, resent on every request, at full price. What you pay frontier rates for that doesn’t need them. Classification, extraction, renaming, formatting. Fix them in that order. The first is free money, the second is usually a 5–10x reduction on the same model, and only then does model choice matter. Here’s what each one is actually worth right now, with numbers from the providers’ own docs rather than from memory. Tier 0: what’s genuinely free today Free tiers are real, and they’re bigger than most people assume. But the numbers that matter aren’t the daily totals everyone quotes — they’re the per-minute ones, because th...