Chuyển đến nội dung chính

Bài đăng

Hiển thị các bài đăng có nhãn Observability

Mobile observability: Sentry, Crashlytics, and custom traces — hướng dẫn Flutter

Crashes without release context are archaeology. Observability is product hygiene. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Correlate crashes with performance and releases. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Tag traces with user flows and feature flags. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure có ki...

Sentry Flutter: errors with release context

Sentry Flutter: errors with release context sentry_flutter is useful because it focuses on crash reporting, performance traces, breadcrumbs, and source maps. The important engineering move is to place the package behind a clear boundary, so your Flutter UI depends on a stable capability rather than a vendor-shaped API. What the library should own Keep package calls inside a named adapter or feature boundary. Make lifecycle, errors, and loading state part of a testable contract. Expose only the capability the app needs; hide implementation details from the whole tree. A focused starting point await SentryFlutter . init ( (options) { options.dsn = const String . fromEnvironment ( 'SENTRY_DSN' ); options.tracesSampleRate = 0.1 ; options.environment = const String . fromEnvironment ( 'APP_ENV' ); }, appRunner : () => runApp ( const App ()), ); Production checklist Read the README and changelog for the exact version you pin....

Mobile observability: Sentry, Crashlytics, and custom traces

Crashes without release context are archaeology. Observability is product hygiene. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Correlate crashes with performance and releases. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Tag traces with user flows and feature flags. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates the int...

Tracing LLM calls inside mobile apps — hướng dẫn Flutter

You cannot improve what you do not measure — including AI latency. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Spans, token metrics, and privacy-safe logging. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Trace request → model → parse → UI paint. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure có kiểu, không trả string ...

Tracing LLM calls inside mobile apps

You cannot improve what you do not measure — including AI latency. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Spans, token metrics, and privacy-safe logging. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Trace request → model → parse → UI paint. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates the intent and calls the bou...

Tracing LLM calls inside mobile apps — hướng dẫn Flutter

You cannot improve what you do not measure — including AI latency. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Spans, token metrics, and privacy-safe logging. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Trace request → model → parse → UI paint. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure có kiểu, không trả string ...

Tracing LLM calls inside mobile apps

You cannot improve what you do not measure — including AI latency. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Spans, token metrics, and privacy-safe logging. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Trace request → model → parse → UI paint. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates the intent and calls the bou...

Observability for LLM apps: tracing a non-deterministic system

“A customer says the assistant told them we offer a 90-day return window. We don’t.” In a normal service you find the request, replay it, and read the code path. In an LLM application, replaying gives you a different answer, the retrieval may have changed, and the code path was identical for the thousand requests that behaved correctly. The trace is not a debugging aid here; it is the only record that the event happened at all. Record the whole run, not the endpoint A single user message can produce a dozen model calls, retrievals and tool invocations. Logging only the final response tells you what went wrong and nothing about where. Model the run as a trace with nested spans: trace: conversation_turn user_id, conversation_id, turn_index ├── span: retrieve query, k, latency, chunk_ids, scores ├── span: llm_call model, temperature, tokens_in/out, stop_reason │ └── span: tool.search_orders arguments, result_size, error, latency ├── span:...

Observability for LLM apps: tracing a non-deterministic system

“A customer says the assistant told them we offer a 90-day return window. We don’t.” In a normal service you find the request, replay it, and read the code path. In an LLM application, replaying gives you a different answer, the retrieval may have changed, and the code path was identical for the thousand requests that behaved correctly. The trace is not a debugging aid here; it is the only record that the event happened at all. Record the whole run, not the endpoint A single user message can produce a dozen model calls, retrievals and tool invocations. Logging only the final response tells you what went wrong and nothing about where. Model the run as a trace with nested spans: trace: conversation_turn user_id, conversation_id, turn_index ├── span: retrieve query, k, latency, chunk_ids, scores ├── span: llm_call model, temperature, tokens_in/out, stop_reason │ └── span: tool.search_orders arguments, result_size, error, latency ├── span:...

Chi phí mỗi request: con số mà hầu hết feature AI không bao giờ tính

Đội nào ship feature AI cũng biết hóa đơn API hàng tháng của mình. Gần như không đội nào biết một request đơn lẻ tốn bao nhiêu. Đó là hai con số khác nhau, và chỉ con số thứ hai mới hành động được. Hóa đơn là một tổng. Nó không nói cho bạn biết tổng đó lớn vì mỗi request đắt, hay vì số request nhiều hơn nhiều so với dự tính. Hai vấn đề này có cách sửa ngược nhau — một cái là vấn đề kiến trúc, một cái là vấn đề đóng gói sản phẩm — nên đội chỉ có mỗi cái tổng thường phản ứng y hệt nhau trong cả hai trường hợp: ai đó bỏ một tuần cắt gọt system prompt. Như bạn sẽ thấy bên dưới, đó thường là đòn bẩy nhỏ nhất hiện có, và là đòn bẩy đầu tiên mọi người với tay tới. Hóa đơn còn là một con số duy nhất cho một tài khoản có thể đang phục vụ sáu feature. Nếu summarization rẻ còn workflow agentic mới thì đắt, hóa đơn gộp chúng thành một con số vô nghĩa và con số đó cứ tăng. Bài này dựng con số còn thiếu đó từ đầu: một hành động người dùng tốn bao nhiêu, từ đầu đến cuối, tính cả mọi hệ số nằm giữa...

Cost per request: the number most AI features never compute

Every team shipping an AI feature knows their monthly API bill. Almost none of them know what a single request costs. Those are different numbers, and only the second one is actionable. The bill is a total. It cannot tell you whether the total is large because each request is expensive or because there are far more requests than you planned for. Those two problems have opposite fixes — one is an architecture problem, the other is a packaging problem — so a team that only has the total tends to respond the same way regardless: someone spends a week trimming the system prompt. As you’ll see below, that is usually the smallest lever available, and it’s the first one everyone reaches for. The bill is also a single scalar for an account that might serve six features. If summarization is cheap and the new agentic workflow is expensive, the invoice averages them into one meaningless number that goes up. What follows builds the missing number from scratch: what one user action costs, end to...