Data Pipeline & Dataset Monitoring
The redesigned flow enabled data engineers to set up alerts independently - removing the reliance on customer success support that had been a consistent blocker in the original experience.
Introduced the ability to set alerts across multiple datasets and pipelines in a single action - a workflow that previously required creating alerts one by one.
Added proactive duplicate detection to the creation flow, reducing alert spam and the CS tickets it generated. Users could now see if a similar alert already existed before creating a new one.
Introduced competitor analysis and iterative user feedback into a sprint cycle that had neither before. The process continued after I joined - a systemic change, not a one-off project.
Databand (acquired by IBM) is a proactive data observability tool for data engineers that helps detect and resolve data issues. The main value Databand gives to its customers is catching bad data before it impacts the business.
Databand's alert system is the core engine that customers use to keep track of their moving data. This is why the alert definition editor is a very strategic and important part of the Databand system. By setting alert rules, customers can choose which metric and what data or pipelines processes they want to monitor.
When I came to Databand, the UX of creating an alert was one of the most painful experiences, and at the same time, one of the most challenging products.
I was the lead designer for the alert creation experience throughout my time at Databand, owning the full design process across 7 iterative releases. This wasn't a single sprint handoff - it was a sustained design effort that evolved as user needs, technical capabilities, and business priorities changed over time.
I ran multi-channel user research (customer support analysis, user interviews, Smartlook session recordings, and demo observations), defined the UX strategy and requirements with the PM, created concepts from lo-fi wireframes through pixel-perfect specs, ran usability tests with real customers, and collaborated closely with the engineering team throughout development and QA. I also introduced the competitor research and iterative feedback practice into the team's sprint cycle - a process that didn't exist before I joined.
Learning how our users were experiencing the old alert creation editor was through 4 main channels: customer support, user interviews, watching recordings of demos for potential users, and watching user recordings on Smartlook.
In order to understand how to design the best and optimal experience for setting an alert on data and pipelines, I was researching direct and indirect competitors, trying to understand how others are answering similar challenges that Databand's users are facing, what kind of possible features we can use, and looking for UX/UI solutions and patterns that can be used as reference or inspiration.
The goal was to design an agile experience that can grow according to the complexity and additional alerts over time.
We knew that for the first iteration we would add a Data Delay alert, so I used this scenario to create a low-fidelity wireframe flow to brainstorm and share ideas inside the product team and get feedback from customer success.
The new alert creation started with a request that kept returning from users at that time. They wanted to know when something is wrong with their data quality. Users wanted to be notified whenever there is an issue with data inside a specific dataset. For that, we needed to add an advanced alert monitor, checking the data source or pipeline process which is relevant.
A possible indication of problematic data is when there is an issue with any of the metadata metrics like data freshness, null count, anomalies, and data statistics. A problematic data might be ingested into the target dataset undetected, and in turn affect downstream data products.
Talking to users, we realized that knowing when data is late to be updated gives them a lot of value in indicating that something is wrong. For that we decided to allow users to set a Data Delay alert across multiple datasets, including the option to target datasets related to a specific pipeline.
To begin with, we started by creating only the alert define experience without a receiver and no alerts gallery. We also temporarily left the old experience under pipeline alert. I created a prototype that would allow us to test the new alert experience with a few customers. It was pretty clear that the new design is easier to consume and much more clear than the old one.
I created detailed specs, presenting different use cases and flows, with high-fidelity design. Together with the front-end team, we defined different interactions and validation behaviors.
Eventually, when the first iteration was on the air, we started collecting feedback from users.
I decided to address the feedback during the following iterations.
According to users' requests, the next iteration included the capability to set alerts on different metrics for column-level data. The new alert included setting alerts to one or more columns of a dataset, on different metrics like anomaly, nulls, and metadata statistics.
When users started working with the new data quality alert, we learned that users are missing 2 important things, that were scoped out because of technical priority:
Those requests were quite critical for users, so we invested the following sprint in adding those requests.
Now was the time to approach another step from the original concept - assigning different receivers for different alerts. Until this point, it was possible to set only one receiver for triggered alerts. The ability to choose different receivers per alert rule was a deal-breaker. Different people inside the same team or company were responsible for different pipelines, and they wanted to get only the triggered alerts relevant to them.
We took advantage to build this feature upon a request that came from an important client; adding integration for the PagerDuty receiver.
The experience was supposed to be very short and simple:
We quickly learned that the UI for selecting a receiver was really hard to understand, which ended up with customers not using it. We updated the experience into a list that exposes all possible receivers from the start, so the user only needs to choose which receiver they want notified.
After introducing the ability to set a data quality alert on multiple datasets, customers asked for the same capability for the 3 most popular pipeline alerts: Run State, Run Duration & Schema Change.
We were looking for a quick way to validate and were wondering if to add multiple selections to the old alert editor, but we faced huge product and UX challenges that solving, would at least double our scope; the old editor supported setting an alert on specific tasks inside pipelines.
We knew that on our future road map we plan to solve the complexity of setting alerts on entire pipelines vs inner tasks. So for this iteration, we decided to add multi selections for those 3 alerts, using the new experience pattern. This would give us a quick solution that we can validate with users.
Our customer success team started to notice that for some reason customers are creating duplicated alerts. talking to them raised a few problems:
We realized that the problem was in our experience, which allowed creating similar alerts without warning the users that they are about to do so.
An alert considers duplicated if there is a similarity in 3 parameters:
In case we recognize that the user setting an alert that is similar to existing, we show him as part of the create alert flow, a validation screen which warn him that he is about to create a duplicated alert, and present him details about those alerts.
Lately, we working on unifying the entire alert experience, and tackling more of the major problems in alert creation products.
I create a low-fidelity mockups in order for the product manager and I are could explore and refine different ideas, and use cases, and collect feedback from customers. Each version starts with assumptions that we validate with customers.
You can see a prototype example for one of the latest concepts.
Note: This version is still in the ideation phase.
We tested this version with our customers, and we learned some good & bad insights.
After the testing we did, we understood the we need to update a little our assumptions:
Over 7 releases, the alert creation experience went from one of Databand's most painful interactions to a self-serve workflow that data engineers could complete without customer success involvement. Key outcomes:
The conversational UX pattern was the right call for simple, single-asset alerts - users responded to it immediately and it tested well. But as the product grew to support multi-asset scenarios, pipeline-vs-task hierarchy, and per-alert receivers, the pattern started showing its limits. We ended up with three different experiences that each solved a piece of the problem but didn't feel like a coherent system.
If I were starting from scratch, I would have invested more time upfront in designing the mental model rather than the interface. The core question - "is this alert about a pipeline, a task inside a pipeline, or a dataset?" - was something users consistently struggled with across every iteration. Getting a cleaner answer to that question early would have shaped a more extensible foundation and saved several rounds of rework.