How do you determine the appropriate watermark threshold for different data sources, and how do you communicate these decisions to your team, especially when choosing a different watermark threshold for new data sources?
💡 Model Answer
Determining a watermark threshold starts with understanding each source’s latency profile and business tolerance for late data. For sources with low, predictable latency (e.g., internal logs), a small watermark lag (e.g., 5–10 seconds) suffices. For external feeds with variable network delays, a larger lag (e.g., 30–60 seconds) or a percentile‑based approach is safer. I gather historical lateness statistics, plot the lateness distribution, and set the watermark to a percentile that balances completeness and latency. I also factor in downstream state TTLs and query semantics. Once thresholds are defined, I document them in a shared runbook, create a dashboard that visualizes current watermark positions per source, and schedule a brief sync with the data engineering team to review the rationale. For new sources, I repeat the profiling step, update the runbook, and communicate changes via a Slack channel or a short meeting, ensuring that all stakeholders understand the impact on downstream analytics and alerting.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500