Chuyển đến nội dung chính

Bài đăng

Hiển thị các bài đăng có nhãn Testing

Eval harnesses for mobile AI features — hướng dẫn Flutter

AI quality is stochastic. Evals make it shippable. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Golden prompts, regression sets, and release gates. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Keep a versioned prompt + dataset in the repo. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure có kiểu, không trả string chỉ dàn...

Eval harnesses for mobile AI features

AI quality is stochastic. Evals make it shippable. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Golden prompts, regression sets, and release gates. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Keep a versioned prompt + dataset in the repo. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates the intent and calls the boundary. ...

Patrol: integration tests that can leave the Flutter sandbox

Patrol: integration tests that can leave the Flutter sandbox patrol is useful because it focuses on native permissions, notifications, deep links, and device automation. The important engineering move is to place the package behind a clear boundary, so your Flutter UI depends on a stable capability rather than a vendor-shaped API. What the library should own Keep package calls inside a named adapter or feature boundary. Make lifecycle, errors, and loading state part of a testable contract. Expose only the capability the app needs; hide implementation details from the whole tree. A focused starting point patrolTest ( 'camera permission flow' , ($) async { await $. native . grantPermissionWhenInUse (); await $. pumpWidgetAndSettle ( const CameraApp ()); await $(#takePhoto). tap (); expect (find. byType ( PhotoPreview ), findsOneWidget); }); Production checklist Read the README and changelog for the exact version you pin. Add one test for lifecycle...

Mocktail: readable Dart mocks without generated files

Mocktail: readable Dart mocks without generated files mocktail is useful because it focuses on fakes, stubbing, verification, fallback values, and test boundaries. The important engineering move is to place the package behind a clear boundary, so your Flutter UI depends on a stable capability rather than a vendor-shaped API. What the library should own Keep package calls inside a named adapter or feature boundary. Make lifecycle, errors, and loading state part of a testable contract. Expose only the capability the app needs; hide implementation details from the whole tree. A focused starting point class MockUserApi extends Mock implements UserApi {} final api = MockUserApi (); when (() => api. fetchUser ( '42' )). thenAnswer ((_) async => user); final result = await repository. load ( '42' ); verify (() => api. fetchUser ( '42' )). called ( 1 ); Production checklist Read the README and changelog for the exact version yo...

Golden Toolkit: visual regression with intent

Golden Toolkit: visual regression with intent golden_toolkit is useful because it focuses on device matrices, fonts, surface variants, and readable golden diffs. The important engineering move is to place the package behind a clear boundary, so your Flutter UI depends on a stable capability rather than a vendor-shaped API. What the library should own Keep package calls inside a named adapter or feature boundary. Make lifecycle, errors, and loading state part of a testable contract. Expose only the capability the app needs; hide implementation details from the whole tree. A focused starting point testGoldens ( 'profile variants' , (tester) async { await loadAppFonts (); await tester. pumpWidgetBuilder ( const ProfileCard ()); await screenMatchesGolden (tester, 'profile-card' ); }); Production checklist Read the README and changelog for the exact version you pin. Add one test for lifecycle, failure, and app background/foreground behavior. Ve...

Pseudolocales and expansion testing before translators — hướng dẫn Flutter

German and Vietnamese expand. Arabic flips. Test before paying translators. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Finding truncation bugs without a full translation set. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Pseudolocales catch overflow for free. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure có kiểu, khô...

Integration tests that survive real devices — hướng dẫn Flutter

Widget tests cannot open system dialogs. Patrol drives native UI alongside Flutter. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi patrol, native automation, and CI device farms for Flutter. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Keep e2e to golden paths; put the rest in widget/unit tests. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data...

Golden tests for design systems

Design systems need visual regression more than feature apps. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Font loading, surface sizes, and review workflow for visual diffs. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Load real fonts in tests or goldens will lie. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates the intent...

Eval harnesses for mobile AI features — hướng dẫn Flutter

AI quality is stochastic. Evals make it shippable. Bài này chuyển quan sát đó thành một mô hình nhỏ để bạn có thể kiểm thử, đo lường và giữ cho code dễ bảo trì khi app lớn lên. Vấn đề cốt lõi Golden prompts, regression sets, and release gates. Câu hỏi hữu ích không phải API hay pattern có đẹp riêng lẻ hay không, mà là state nằm ở đâu, boundary nào chịu trách nhiệm khi lỗi xảy ra, và người dùng phục hồi thế nào khi happy path biến mất. Mô hình thực tế Keep a versioned prompt + dataset in the repo. Hãy bắt đầu bằng một owner rõ ràng cho behavior. Widget chỉ nên render và phát intent; IO, persistence, permission và retry nên nằm sau một interface nhỏ. Nhờ vậy bạn có seam để fake trong test và một chỗ ghi lại các dữ kiện cần theo dõi ở production. Trong app Flutter, boundary thường có dạng: Widget phát intent như load, submit, refresh hoặc retry. Controller hoặc use case validate intent rồi gọi boundary. Boundary trả về data có kiểu hoặc failure có kiểu, không trả string chỉ dàn...

Tiêm phụ thuộc trong Flutter mà không cần nghi thức rườm rà

Tiêm phụ thuộc có một cái tên đáng sợ cho một ý tưởng hết sức bình thường: một lớp nên được trao thứ nó cần thay vì tự dựng hoặc tự đi tìm. Toàn bộ khái niệm chỉ có vậy. Mọi thứ còn lại — container, locator, provider, mã sinh tự động — chỉ là bộ máy để giao hàng. Lý do đáng quan tâm là kiểm thử. Viết ApiClient() bên trong một repository nghĩa là mọi test của repository đó đều gọi mạng thật. Truyền ApiClient vào nghĩa là mọi test đều truyền được bản giả. Đó là toàn bộ phần thưởng, và thế là đủ. Tiêm qua constructor, không cần gói nào final class UserRepository { const UserRepository ( this ._api, this ._cache); final ApiClient _api; final UserCache _cache; Future < User > fetch ( String id) async { final cached = _cache. get (id); if (cached != null ) return cached; final user = await _api. getUser (id); _cache. put (user); return user; } } Kiểm thử nó không cần framework nào: test ( 'trả về user từ cache mà...

Một pipeline CI cho Flutter thật sự bắt được lỗi

Kho mã Flutter nào rồi cũng mọc ra một file .github/workflows/ci.yml chứa flutter test . Nó pass, mọi người thấy yên tâm, rồi một bản build phát hành hỏng trên máy build vì lý do mà không laptop nào có thể lộ ra. CI đáng đồng tiền khi nó bắt được những lỗi mà môi trường phát triển cục bộ về mặt cấu trúc không thể bắt: mã sinh cũ, một phụ thuộc chỉ resolve được nhờ pub cache trên máy bạn, một thay đổi định dạng chưa ai chạy, một golden đã trôi. Đây là pipeline dựng quanh những thứ đó. Hình dạng tổng thể Bốn job, chia hai đợt: Job Chạy khi Bắt được analyze mọi push và PR lệch định dạng, hồi quy lint, mã không dùng test mọi push và PR lỗi unit, widget và golden build PR vào main và tag lỗi biên dịch chỉ xảy ra trên máy sạch release chỉ tag ký, tải artifact lên analyze và test chạy song song và nhanh. build chậm nên bị chặn cổng. Sự phân chia đó quan trọng: một pipeline mà mỗi lần push đều phải chờ tám phút build Android là pipeline mà người ta sẽ học...

Pseudolocales and expansion testing before translators

German and Vietnamese expand. Arabic flips. Test before paying translators. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Finding truncation bugs without a full translation set. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Pseudolocales catch overflow for free. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates the intent and...

Integration tests that survive real devices

Widget tests cannot open system dialogs. Patrol drives native UI alongside Flutter. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence patrol, native automation, and CI device farms for Flutter. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Keep e2e to golden paths; put the rest in widget/unit tests. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or...

Eval harnesses for mobile AI features

AI quality is stochastic. Evals make it shippable. This article turns that observation into a small implementation model you can test, measure, and keep boring when the app grows. The problem in one sentence Golden prompts, regression sets, and release gates. The useful question is not whether the API or pattern looks elegant in isolation. It is where state lives, which boundary owns failure, and how a user recovers when the happy path disappears. A practical model Keep a versioned prompt + dataset in the repo. Start with one explicit owner for the behavior. Keep widgets responsible for rendering and user intent; keep IO, persistence, permissions, and retries behind a small interface. That gives you a seam for a fake in tests and a place to record the facts that matter in production. For a Flutter app, the boundary usually looks like this: The widget emits an intent such as load, submit, refresh, or retry. A controller or use case validates the intent and calls the boundary. ...

Dữ liệu tổng hợp cho eval: xây bộ kiểm thử mà bạn tin được

Bạn có một hệ thống RAG, chưa có người dùng, và một prompt cứ sửa đi sửa lại. Lần sửa nào cũng thấy “có vẻ tốt hơn” mà không cách nào kiểm chứng. Viết tay một trăm câu hỏi kiểm thử mất hai ngày, và bạn sẽ không làm. Thế là bạn nhờ model sinh ra chúng. Cách này có tác dụng, đáng làm, và có đúng một kiểu hỏng khiến toàn bộ nỗ lực trở nên vô nghĩa nếu bạn không thiết kế để né nó. Vấn đề luẩn quẩn Nếu bạn sinh câu hỏi bằng cách đưa tài liệu cho model đọc, bạn sẽ nhận về những câu hỏi mà tài liệu trả lời tốt, diễn đạt theo đúng cách tài liệu diễn đạt. Retrieval của bạn sẽ đạt điểm rất đẹp — trên một bộ test được dựng từ chính những giả định của hệ thống đang bị kiểm tra. Người dùng thật thì hỏi về những thứ tài liệu bao phủ kém, bằng từ ngữ tài liệu chưa bao giờ dùng, kèm theo tiền đề sai. Một bộ eval tổng hợp làm ngây thơ chỉ đo tính nhất quán nội bộ, không đo tính hữu ích. Nó sẽ đạt trong khi người dùng của bạn thất bại. Toàn bộ phần dưới đây là về việc bẻ gãy vòng luẩn quẩn đó. Gi...