The full question
When a client app loads, it needs to fetch everything required to render the first screen in a single call. That data lives behind three separate internal services, so you will build an aggregator (a "bootstrap" endpoint) that fans out to them and composes one unified response.
Downstream services
You are given three internal services (internal APIs):
- User Service —
GET /user-to-consumer?user_id=... → returns { consumer_id, user_profile... } - Payments Service —
GET /payment-info?consumer_id=... → returns { payment_methods... } - Address Service —
GET /address-info?consumer_id=... → returns { addresses... }
Note the dependency chain: the Payments and Address services are keyed on consumer_id, which only the User Service can produce from a user_id.
What to build
Design and implement a Bootstrap API:
- Endpoint:
GET /bootstrap?user_id=... - Behavior: take the input
user_id, fetch the corresponding data from the downstream services, and return a single response that aggregates: - user / profile information
- payment information
- address information
Core requirement
The endpoint must be as resilient to failures as possible. Downstream services may be slow, timing out, erroring, or intermittently / partially unavailable, and the bootstrap response should degrade gracefully rather than fail outright.
---
Constraints & Assumptions
Anchor your design with the following working assumptions (confirm or adjust them with the interviewer):
- The endpoint is on the client's first-paint critical path, so it is latency-sensitive — assume a target such as p99 ≤ ~600 ms end-to-end.
- The operation is a read-only
GET (inherently idempotent). - Typical microservice constraints apply: bounded thread/connection pools, shared infrastructure, no distributed transactions.
- Treat exact SLO numbers, retry counts, and TTLs as tunable — state the figures you choose rather than leaving them implicit.
---
Part 1 — The API Contract
Define the response shape and, crucially, what happens on partial failures — when some downstream data is available and some is not. Specify the HTTP-level and body-level semantics a client can program against.
Part 2 — Orchestration: Ordering & Concurrency
Describe how you sequence and parallelize the three calls, and how you bound total latency.
Part 3 — Reliability Strategies
Cover timeouts, retries, circuit breakers, fallbacks, and caching, and how they compose.
Part 4 — Observability & Operational Considerations
Describe what you measure, alert on, and can tune at runtime.
---
Clarifying Questions to Ask
A strong candidate scopes the problem before designing. Reasonable questions include:
- How does each downstream affect what the screen can render? Can the response still be useful if one or two sections are missing, or does the client treat all three as mandatory? Which call, if any, blocks every other piece of the response?
- Who calls this and how is it authenticated? Is
user_id trusted from the query string, or must it be derived from the authenticated principal? - What is the latency SLO for first paint, and what is the per-call budget for each downstream?
- Are partial responses acceptable to the client, or must all three sections be present atomically?
- What is the staleness tolerance per section — can addresses / profile / payment methods be served from cache, and for how long?
- What is the read volume / fan-out scale, and are there per-user rate limits to respect?
- Are there correctness constraints on payments specifically (e.g. must we never display a removed or expired payment method)?
Model answer
### Part 1 — The API Contract
**Response Shape:**
The response from the `/bootstrap?user_id=...` endpoint will have the following structure:
```json
{
"user_profile": {
"consumer_id": "12345",
"name": "John Doe",
"email": "john.doe@example.com"
},
"payment_methods": [
{
"method_id": "1",
"type": "credit_card",
"last4": "1234"
},
{
"method_id": "2",
"type": "paypal"
}
],
"addresses": [
{
"address_id": "1",
"street": "123 Main St",
"city": "Anytown",
"state": "CA",
"zip": "90210"
}
],
"errors": {
"user_service": null,
"payment_service": "Timeout fetching payment methods.",
"address_service": null
}
}
Partial Failures:
- If a downstream service fails, the corresponding section will be
null, and an error message will be logged in the errors object. - The response will still include available data, allowing the client to render as much as possible.
Part 2 — Orchestration: Ordering & Concurrency
- Fetch User Profile: - Call
GET /user-to-consumer?user_id=... to retrieve consumer_id and user profile. - Fetch Payment and Address Info Concurrently: - Once
consumer_id is received, make two concurrent calls: - GET /payment-info?consumer_id=... - GET /address-info?consumer_id=... - Total Latency Bound: - Set a timeout of 200 ms for each downstream call. - Use a total timeout of 600 ms for the entire bootstrap call. - If any call exceeds its timeout, handle it gracefully and log the error.
Part 3 — Reliability Strategies
- Timeouts:
- Set individual timeouts of 200 ms for each downstream service call.
- Retries:
- Implement a maximum of 1 retry for each service call, with exponential backoff.
- Circuit Breakers:
- Use a circuit breaker pattern to prevent overwhelming a failing service.
- If a service fails consecutively, open the circuit for a defined period (e.g., 30 seconds).
- Fallbacks:
- If the payment or address service fails, return cached data if available.
- Caching:
- Cache user profiles for 5 minutes, payment methods for 1 minute, and addresses for 10 minutes.
Part 4 — Observability & Operational Considerations
- Metrics to Measure:
- Response time for the
/bootstrap endpoint. - Success/failure rates for each downstream service.
- Latency for each individual service call.
- Alerts:
- Set alerts for high failure rates (>5% error rate) and latency spikes (>500 ms).
- Runtime Tuning:
- Allow adjustment of timeout values, retry counts, and cache TTLs via configuration.