正在加载内容...

963963 Chat Topics Portal Independent coverage of news

Common Mistakes When Evaluating Observability

By James Whitfield · · 1222 words
Common Mistakes When Evaluating Observability

Consider queue design specifically. The interesting number is not the average, it is the 99th percentile. Queue Design: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to queue design as well.

Storage Tiers: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. That applies to storage tiers as well. In practice, storage tiers behaves differently: Failures are usually correlated, so plan for the shared dependency.

You can often replace a coordination problem with an idempotency key. That applies to observability as well. In practice, observability behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for observability.

Queue Design: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to queue design as well. In practice, queue design behaves differently: Costs usually concentrate in a small number of operations, so find those first.

Configurations should be reviewable in a diff, not only in a console. This is most visible in crawl budget. Consider crawl budget specifically. The best time to add an index is before the table gets large. Crawl Budget: Failures are usually correlated, so plan for the shared dependency.

Use statements about your own needs rather than trying to guess your partner’s intentions. You might say, “I’m comfortable with this, but not with that,” or, “I need us to stop if I say pause.” Be specific about what you mean by words such as “slow down” or “check in.” Ask your partner what they are comfortable with, and leave room for an answer without interrupting or arguing.

Load Balancing: You can often replace a coordination problem with an idempotency key. Load Balancing: Anything that grows without a bound will eventually hit one. Load Balancing: Documentation that is not tested tends to describe the previous version.

Crawl Budget: You can often replace a coordination problem with an idempotency key. Crawl Budget: Anything that grows without a bound will eventually hit one. Crawl Budget: Documentation that is not tested tends to describe the previous version.

For edge caching, the constraint matters more than the feature list. The first thing to settle is the failure mode, not the happy path. Teams working on edge caching usually discover this the hard way. Measurements taken once are anecdotes; you need a baseline that repeats. Costs usually concentrate in a small number of operations, so find those first. This is most visible in edge caching.

Teams working on search indexing usually discover this the hard way. Serving static bytes is the cheapest thing you can do at the edge. A schema is an interface; changing it is a migration, not an edit. This is most visible in search indexing. Consider search indexing specifically. Track the denominator as carefully as the numerator.

Content Delivery: A queue smooths spikes but also hides how far behind you are. Content Delivery: Retries without jitter turn a small outage into a large one. Content Delivery: Separating the reads from the writes buys room to change either side.

Monitoring Alerts: The first thing to settle is the failure mode, not the happy path. Monitoring Alerts: Measurements taken once are anecdotes; you need a baseline that repeats. Monitoring Alerts: Costs usually concentrate in a small number of operations, so find those first.

Teams working on access control usually discover this the hard way. You can often replace a coordination problem with an idempotency key. Anything that grows without a bound will eventually hit one. This is most visible in access control. Consider access control specifically. Documentation that is not tested tends to describe the previous version.

Cloud Infrastructure: A design that cannot be rolled back is a design that cannot be changed safely. Cloud Infrastructure: Latency budgets are easier to defend when every hop has a stated ceiling. Cloud Infrastructure: Caching helps only until the invalidation rules become the bottleneck.

In practice, load balancing behaves differently: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Aggregating at write time trades flexibility for predictable read cost.

Data Pipelines: The first thing to settle is the failure mode, not the happy path. Data Pipelines: Measurements taken once are anecdotes; you need a baseline that repeats. Data Pipelines: Costs usually concentrate in a small number of operations, so find those first.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on load balancing usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

People may communicate boundaries differently, and no single gesture reliably proves consent. Look for clear, freely given agreement, but do not rely on body language alone when you are unsure. If communication is difficult, slow down and agree on words or signals that both people understand before continuing.

Consider api design specifically. If the rollback plan needs a meeting, it is not a rollback plan. API Design: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. That applies to api design as well.

Monitoring Alerts: If the rollback plan needs a meeting, it is not a rollback plan. Monitoring Alerts: Small pages that stay small are easier to keep fast than large ones made fast. Monitoring Alerts: Write the invariant down; otherwise it lives only in someone's memory.

Access Control: Serving static bytes is the cheapest thing you can do at the edge. Access Control: A schema is an interface; changing it is a migration, not an edit. Access Control: Track the denominator as carefully as the numerator.

A design that cannot be rolled back is a design that cannot be changed safely. The same reasoning holds for monitoring alerts. For monitoring alerts, the constraint matters more than the feature list. Latency budgets are easier to defend when every hop has a stated ceiling. Teams working on monitoring alerts usually discover this the hard way. Caching helps only until the invalidation rules become the bottleneck.

You can also state your own boundaries. Say what you are comfortable with and what you do not want, and ask questions if an answer is unclear. Good communication is not a guarantee that everything will go as expected; it is a way to make choices more explicit and respond when circumstances change.

Rate Limiting: Configurations should be reviewable in a diff, not only in a console. Rate Limiting: The best time to add an index is before the table gets large. Rate Limiting: Failures are usually correlated, so plan for the shared dependency.

Related reading