# Welcome!

Free Open-source ML observability course for data scientists and ML engineers by Evidently AI.

## Welcome!

{% embed url="<https://www.youtube.com/watch?v=FI2xDByfXfQ&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=2>" %}

\
[ML observability course: welcome video](https://www.youtube.com/watch?v=FI2xDByfXfQ\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=2)

Welcome to the Open-source ML observability course!

## How to participate?

* **Learn at your own pace**. We published all 40 lessons with videos, course notes, and code examples.
* **Join the course cohort**. To submit assignments and earn a certificate of completion, you must enroll in the course cohort. [Sign up](https://www.evidentlyai.com/ml-observability-course) to save your seat and be notified when the next cohort starts.

## Links

* **Newsletter**. [Sign up](https://www.evidentlyai.com/ml-observability-course) to receive course updates and be notified when the next cohort starts.
* **Discord community**. Join the [community](https://discord.gg/PyAJuUD5mB) to ask questions and chat with others.
* **Code examples**. Are published in this GitHub [repository](https://github.com/evidentlyai/ml_observability_course).
* **YouTube playlist**. [Subscribe](https://www.youtube.com/playlist?list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF) to the course YouTube playlist.

**Enjoying the course?** [Star](https://github.com/evidentlyai/evidently) Evidently on GitHub to contribute back! This helps us create free, open-source tools and content for the community.

## What the course is about

This course is a deep dive into ML model observability and monitoring.

We explore different types of evaluations, from data quality to data drift, and how this fits in the model lifecycle. We also cover the engineering aspect of ML observability and how to integrate it with your ML services and pipelines.

## Course structure

ML observability course is organized into six modules. You can follow the complete course syllabus or pick only the modules that are most relevant to you.

{% content-ref url="/pages/lcl7lXIR2Da2OWmaragk" %}
[Module 1: Introduction](/ml-observability-course/module-1-introduction)
{% endcontent-ref %}

{% content-ref url="/pages/jcVt55pbTm4BvLCat0SV" %}
[Module 2: ML monitoring metrics](/ml-observability-course/module-2-ml-monitoring-metrics)
{% endcontent-ref %}

{% content-ref url="/pages/ifyhDjTltvJfNvenGnHI" %}
[Module 3: ML monitoring for unstructured data](/ml-observability-course/module-3-ml-monitoring-for-unstructured-data)
{% endcontent-ref %}

{% content-ref url="/pages/Ttu1VHvKy9zXWaXNX4ss" %}
[Module 4: Designing effective ML monitoring](/ml-observability-course/module-4-designing-effective-ml-monitoring)
{% endcontent-ref %}

{% content-ref url="/pages/mF1AJg1HRPBJxoOqM9zK" %}
[Module 5: ML pipelines validation and testing](/ml-observability-course/module-5-ml-pipelines-validation-and-testing)
{% endcontent-ref %}

{% content-ref url="/pages/Q7Ga4uWHB7Oh63wCEEEQ" %}
[Module 6: Deploying an ML monitoring dashboard](/ml-observability-course/module-6-deploying-an-ml-monitoring-dashboard)
{% endcontent-ref %}

## Course calendar and deadlines for the 2023 cohort

The 2023 cohort has completed. You can learn at your own pace or [sign up](https://www.evidentlyai.com/ml-observability-course) for the next cohort.

| Module                                                                                                                                                                       | Week                                                                        |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- |
| [Module 1: Introduction to ML monitoring and observability](https://learn.evidentlyai.com/ml-observability-course/module-1-introduction)                                     | October 16, 2023                                                            |
| [Module 2: ML monitoring metrics: model quality, data quality, data drift](https://learn.evidentlyai.com/ml-observability-course/module-2-ml-monitoring-metrics)             | October 23, 2023                                                            |
| [Module 3: ML monitoring for unstructured data: NLP, LLM and embeddings](https://learn.evidentlyai.com/ml-observability-course/module-3-ml-monitoring-for-unstructured-data) | October 30, 2023                                                            |
| [Module 4: Designing effective ML monitoring](https://learn.evidentlyai.com/ml-observability-course/module-4-designing-effective-ml-monitoring)                              | November 6, 2023                                                            |
| [Module 5: ML pipelines validation and testing](https://learn.evidentlyai.com/ml-observability-course/module-5-ml-pipelines-validation-and-testing)                          | November 13, 2023                                                           |
| [Module 6: Deploying an ML monitoring dashboard](https://learn.evidentlyai.com/ml-observability-course/module-6-deploying-an-ml-monitoring-dashboard)                        | November 20, 2023                                                           |
| Final assignment                                                                                                                                                             | <p>November 27, 2023<br><br>Quizzes and assignment due December 4, 2023</p> |

## Our approach

* **Blend of theory and practice**. The course combines key concepts of ML observability and monitoring with practice-oriented tasks.
* **Practical code examples**. We provide end-to-end deployment blueprints and walk you through the code examples.
* **Focus on open-source**. The course is built upon open-source tools to make ML observability accessible to all.
* **The course is free and open to everyone**. All course videos are public so you can rewatch them anytime.

## Prerequisites

There are both theoretical and code-focused modules that require knowledge of Python. We will walk you through the code, but you can skip these parts and still learn a lot.

## Who is it for

This course is useful to professionals who have dealt with ML models in production and those preparing to deploy ML models:

* Data scientists,
* ML engineers,
* Technical product managers,
* Analysts.

Let’s dive in!


# Module 1: Introduction

Key concepts of machine learning monitoring and observability and how they fit in the ML lifecycle.

This theoretical module introduces the key topics of machine learning monitoring and observability.

It covers the following topics:

* what can go wrong with ML models in production;
* what ML monitoring and observability are and how they fit in the ML lifecycle;
* what types of evaluation you might need, from model quality to data drift;
* key considerations to keep in mind when designing your monitoring.

At the end of this module, you will know the key concepts related to ML monitoring and observability and how they will be covered throughout the course.


# 1.1. ML lifecycle. What can go wrong with ML in production?

What can go wrong with data and machine learning services in production. Data quality issues, data drift, and concept drift.

{% embed url="<https://www.youtube.com/watch?v=8I89FY2eelM&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=2>" %}

**Video 1**. [ML lifecycle. What can go wrong with ML in production](https://www.youtube.com/watch?v=8I89FY2eelM\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=2), by Emeli Dral

## Evaluations in ML model lifecycle

Building a successful ML model involves the following stages:

* Data preparation,
* Feature engineering,
* Model training,
* Model evaluation,
* Model deployment.

You can perform different types of evaluations at each of these stages. For example,

* During data preparation, exploratory data analysis (EDA) helps to understand the dataset and validate the problem statement.
* At the experiment stage, performing cross-validation and holdout testing helps validate and test if ML models are useful.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a2e44d04d0c1404c3f3cd3d15abfc85b312cd947%2F2023109_course_module1_fin_images.005-min.png?alt=media)

However, the work does not stop here! Once the best model is deployed to production and starts bringing business value, every erroneous prediction has its costs. It is crucial to ensure that this model functions stably and reliably. To do that, one must continuously monitor the production ML model and data.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-c198e9d2546bf05dd0052fd139bc84171bc67dcf%2F2023109_course_module1_fin_images.008-min.png?alt=media)

## What can go wrong in production?

Many things can go wrong once you deploy an ML model to the real world. Here are some examples.

**Training-serving skew**. Model degrades if training data is very different from production data.

**Data quality issues**. In most cases, when something is wrong with the model, this is due to data quality and integrity issues. These can be caused by:

* Data processing issues, e.g., broken pipelines or infrastructure updates.
* Data schema changes in the upstream system, third-party APIs, or catalogs.
* Data loss at source when dealing with broken sensors, logging errors, database outages, etc.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a29dc4278a12ee293c88565261ded2ce61304035%2F2023109_course_module1_fin_images.011-min.png?alt=media)

**Broken upstream model**. Often, not one model but a chain of ML models operates in production. If one model gives wrong outputs, it can affect downstream models.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-37c04ae626c0256d547be2bffb05e4d34db44fab%2F2023109_course_module1_fin_images.012-min.png?alt=media)

**Concept drift**. Gradual concept drift occurs when the target function continuously changes over time, leading to model degradation. If the change is sudden – like the recent pandemic – you’re dealing with sudden concept drift.

**Data drift**. Distribution changes in the input features may signal data drift and potentially cause ML model performance degradation. For example, a significant number of users coming from a new acquisition channel can negatively affect the model trained on user data. Chances are that users from different channels behave differently. To get back on track, the model needs to learn new patterns.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-72df0b982113d0de327c06f817c71c291c5bb2ae%2F2023109_course_module1_fin_images.015-min.png?alt=media)

**Underperforming segments**. A model might perform differently on diverse data segments. It is crucial to monitor performance across all segments.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-bdf3d9a1e0e83e0ad4ee433bb3794b49e3adc18c%2F2023109_course_module1_fin_images.016-min.png?alt=media)

**Adversarial adaptation**. In the era of neural networks, models might face adversarial attacks. Monitoring helps detect these issues on time.

## Summing up

Many factors can impact the performance of an ML model in production. ML monitoring and observability are crucial to ensure that models perform as expected and provide value.


# 1.2. What is ML monitoring and observability?

What ML monitoring is, the challenges of production ML monitoring, and how it differs from ML observability.

{% embed url="<https://www.youtube.com/watch?v=Wfphz6TUikM&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=3>" %}

**Video 2**. [What is ML monitoring and observability](https://www.youtube.com/watch?v=Wfphz6TUikM\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=3), by Emeli Dral

## What is ML monitoring?

**Machine learning model monitoring** is a series of techniques to track and analyze the performance of ML models in production. It helps measure ongoing model quality, detect potential issues, and resolve them on time.

## When to monitor?

There are three main scenarios when you need ML model monitoring:

* **Models in production**. Upon deploying ML models to production, you need to keep tabs on the ongoing performance and business impact.
* **Models in shadow deployment**. In shadow mode, you track the behavior of a candidate model when predictions are generated but not acted upon (in other words, ML models generate outputs, but these outputs are not used in downstream systems).
* **During A/B testing**. In this case, you track and compare the results of active candidate models.

## Challenges of ML monitoring

**It’s not just software performance**. Software monitoring has been around for ages. If a software service is running in production, you need to monitor service health metrics such as uptime, memory usage, and latency. This also applies to ML monitoring.

**However, ML monitoring also introduces two extra layers to keep tabs on**:

* Data health,
* Model health.

Accordingly, you also need to track data quality and model performance metrics.

**Selecting the right metric and threshold can be a challenging task**. It is always case-specific. Let’s take a binary classification problem as an example:

* For balanced datasets, an accuracy score of 99% means that the model performs pretty well.
* However, for unbalanced datasets (e.g., in fraud detection tasks), the same accuracy score does not allow making meaningful conclusions about the model quality.

**Ground truth is not available immediately** to calculate ML model performance metrics. In this case, you can use proxy metrics like data quality to monitor for early warning signs.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-05c77b659c5c7f9503d6c007f43fdbb2a4fddf80%2F2023109_course_module1_fin_images.024-min.png?alt=media)

## ML monitoring vs ML observability

Often used interchangeably, ML monitoring and ML observability are quite different.

**ML monitoring**:

* tracks a pre-defined set of metrics,
* helps detect issues (“What happened?”, “Is the system working?”),
* is more reactive (helps to find “known unknowns”).

**ML observability**:

* gives visibility into the system behavior,
* helps understand and analyze root causes (“Why it happened? Where exactly?”),
* is more proactive (helps to find “unknown unknowns”).

In this course, we refer to ML monitoring as the subset of ML observability.

## Why ML monitoring and observability matter

ML monitoring and observability help:

* **Detect issues** and alert about missing data, features out of expected range, data drift, or sudden model quality drops.
* **Find a root cause** by locating corrupted data, detecting low-performing segments, helping select the right data to label, etc.
* **Understand ML model behavior**. It provides insights into how users interact with the model or whether there are changes in the environment where the model functions.
* **Trigger actions**. Based on the calculated data and model health metrics, you can trigger fallback, model switching, or automatic retraining.
* **Document ML model performance** to provide information to the stakeholders.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-382b72e231043344dc34b3463a3e5eea6cb37697%2F2023109_course_module1_fin_images.030-min.png?alt=media)

## Who should care about ML monitoring and observability?

The short answer: everyone who cares about the model's impact on business. At the same time, different users might care about specific aspects of the ML model performance. For example:

* **Data scientists** might be interested in model performance and data drift metrics.
* **Data science managers** need to know whether it is time to retrain the model.
* **Product managers** want to understand model limitations.
* **Data engineers** want to ensure data quality and integrity.

Other stakeholders include model users, business stakeholders, support, and compliance teams.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-1825a5bf013055923c8ac245d2cd4e13aef3b68c%2F2023109_course_module1_fin_images.031-min.png?alt=media)

## Summing up

Often used interchangeably, ML monitoring and ML observability are quite different. While ML monitoring helps detect issues, ML observability helps understand and analyze their root causes and proactively explore the model performance. In this course, we’ll cover both and refer to ML monitoring as the subset of ML observability.

There are many use cases where ML monitoring and observability are crucial for the overall model's well-being, from detecting and resolving issues to understanding processes behind the model. ML monitoring might involve different stakeholders, from data scientists and ML engineers to business teams.


# 1.3. ML monitoring metrics. What exactly can you monitor?

A framework to organize ML monitoring metrics. Software system health, data quality, ML model quality, and business KPIs.

{% embed url="<https://www.youtube.com/watch?v=DCFvZvpDks0&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=4>" %}

**Video 3**. [ML monitoring metrics](https://www.youtube.com/watch?v=DCFvZvpDks0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=4), by Emeli Dral

An ML-based service is more than just an ML model. One needs to keep tabs on all the facets of the ML system quality. When it comes to monitoring the system performance, there are different groups of metrics.

## Software system health

It doesn’t matter how excellent your model is when the whole ML system is down. To track the overall system health, you can reuse existing monitoring schemes from other production services. Standard software performance metrics include latency, error rate, memory usage, disk usage, etc.

## Data quality and data integrity

In many cases, model issues stem from issues with the input data. To monitor data quality and integrity, you can keep tabs on metrics like the share of missing values, type mismatch, or range violations for important features. The goal here is to ensure the stability of data pipelines.

## ML model quality and relevance

ML model performance metrics help to ensure that ML models work as expected:

* Standard metrics help evaluate the quality of the ML model in production. For example, you can track metrics like precision and recall for classification, MAE or RMSE for regression, or top-k accuracy for ranking.
* You can also track use-case specific quality metrics like bias or fairness: for example, through metrics like predictive parity or equalized odds.
* When ground truth is unavailable or delayed, use proxy metrics. Keep tabs on prediction drift, input data drift, or share of new categories. These metrics can signal potential problems before the ML model quality is affected.

## Business KPIs

The ultimate measure of the model quality is its impact on the business. Depending on business needs, you may want to monitor clicks, purchases, loan approval rates, cost savings, etc. This is typically custom to the use case and might involve collaborating with product managers or business teams to determine the right business KPIs.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a7546a15b17982ef67fd3868b11369caa30098b7%2F2023109_course_module1_fin_images.034-min.png?alt=media)

For a deeper dive into **ML model quality and relevance** and **data quality and integrity** metrics, head to [Module 2](/ml-observability-course/module-2-ml-monitoring-metrics).


# 1.4. Key considerations for ML monitoring setup

Key considerations for ML monitoring setup. Service criticality, retraining cadence, reference dataset, and ML monitoring architecture.

{% embed url="<https://www.youtube.com/watch?v=LnfL9Nu0tm4&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=5>" %}

**Video 4**. [Key considerations for ML monitoring setup](https://www.youtube.com/watch?v=LnfL9Nu0tm4\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=5), by Emeli Dral

When designing the ML monitoring setup for a specific model, you might want to consider the following aspects:

* Matching the ML monitoring setup to the use case
* Model retraining cadence
* Choice of reference dataset
* Custom metrics

## Matching ML monitoring setup and the use case

While setting up an ML monitoring system, it makes sense to align the complexity of monitoring with the complexity of the deployment and operations of the ML service. Some factors to consider:

* **ML service implementation**. Is it a real-time production service, batch Airflow DAG, or an ad hoc Python script?
* **Feedback loop and environmental stability**. Both influence the cadence of metrics calculations and the choice of specific metrics.
* **Service criticality**. What is the business cost of model quality drops? What risks should we monitor for? More critical models might require a more complex monitoring setup.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-f9eeba20ea09914cb21762ba6ceaa4d6bada4dfc%2F2023109_course_module1_fin_images.050-min.png?alt=media)

## Model retraining cadence

ML monitoring and retraining are closely connected. Some retraining factors to keep in mind when setting up an ML monitoring system include:

* Frequency and costs of model retraining.
* How you implement the retraining: whether you want to monitor the metrics and retrain on a trigger or set up a predefined retraining schedule (for example, weekly).
* Issues that prevent updating the model too often, e.g., complex approval processes, regulations, need for manual testing.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-c78631b4fbeccdaf87457bd117ec933603add5f4%2F2023109_course_module1_fin_images.052-min.png?alt=media)

## Reference dataset

In a situation where various production models use different data types – e.g., numerical and categorical features, tabular data, text data, images, or videos – setting up data quality monitoring can be overwhelming.

Instead, you can use a **reference dataset** to help automatically generate different tests based on the provided example and compare the new batches of data against it.

If you follow this strategy, it is important that you select and curate an appropriate reference: as it becomes as important as choosing the right metrics. A “good” reference dataset must represent the expected data patterns correctly.

You can also utilize a reference dataset as a baseline for the distribution drift comparison. You can consider having a fixed reference dataset, a moving one, or multiple windows.

Based on the scenario, you can use different reference datasets: for example, one dataset for distribution drift detection and another to generate data quality test conditions.

## Custom metrics

Standard monitoring metrics like accuracy or AUC are good starting points. However, depending on the use case, you may need to introduce more comprehensive custom monitoring metrics.

Some examples of custom metrics include:

* **Use-case specific model quality metrics** (e.g., lift-10% for churn prediction in the telecom industry),
* **Heuristics** that reflect quality (e.g., the share of predictions higher than a specific threshold), especially when ground truth is not available.
* **Business quality metrics** and KPIs (e.g., estimated savings),
* **Custom drift detection methods** beyond standard statistical tests.

## Summing up

While designing an ML monitoring system, tailor your approach to fit your specific requirements and challenges:

* Ensure the monitoring setup aligns with the complexity of your use case.
* Consider binding retraining to monitoring, if relevant.
* Use reference datasets to simplify the monitoring process but make sure they are carefully curated.
* Define custom metrics that fit your problem statement and data properties.

For a deeper dive into the ML monitoring setup, head to [Module 4](/ml-observability-course/module-4-designing-effective-ml-monitoring).


# 1.5. ML monitoring architectures

ML monitoring architectures for backend and frontend and what to consider when choosing between them.

{% embed url="<https://www.youtube.com/watch?v=VVO6QFVbwTU&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=6>" %}

**Video 5**. [ML monitoring architectures](https://www.youtube.com/watch?v=VVO6QFVbwTU\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=6), by Emeli Dral

It is essential to start monitoring ML models as soon as you deploy them to production. To do that, you also need to decide how exactly you implement the ML monitoring process. There are various ML monitoring architectures to choose from, both when it comes to the monitoring backend and frontend.

## Monitoring backend

**Batch ML monitoring**. To set up a batch ML monitoring system, create a pipeline with metric calculations and schedule an ML monitoring job. You can run monitoring jobs periodically (e.g., every 10 seconds, hourly) or on a trigger (e.g., when a new batch of labeled data arrives). As some monitoring metrics like data drift are calculated in batch mode, this architecture works for both batch and online inference.

**Near real-time (streaming) ML monitoring**. In this case, you send data from the ML service to an ML monitoring system directly, calculate monitoring metrics on the fly, and visualize them on an online dashboard. This type of architecture is suited for ML models that require immediate or near-real-time monitoring.

**Ad hoc reporting** is a great alternative when your resources are limited. You can use Python scripts to calculate and analyze metrics in your notebook. This is a good first step in logging model performance and data quality.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-c8ec10ab36f649e9b2fac7d6245b6d3748a27e06%2F2023109_course_module1_fin_images.061-min.png?alt=media)

## Monitoring frontend

When it comes to visualizing the results of monitoring, you also have options.

**No user interface**. You can start with collecting the data without visualizing it. In this scenario, you can simply log metrics to a database. This allows you to implement alerting and historical data logging.

**One-off reports**. You can also generate reports as needed and create visualizations or specific one-off analyses based on the model logs. You can create your own reports in Python/R or use different BI visualization tools.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-ef8292f9bf7a7fad7771be260579a0c35d2b3a00%2F2023109_course_module1_fin_images.065-min.png?alt=media)

**BI Systems**. If you want to create a dashboard to track ML monitoring metrics over time, you can also reuse existing business intelligence or software monitoring systems. In this scenario, you must connect existing tools to the ML metric database and add panels or plots to the dashboard.

**Dedicated ML monitoring**. As a more sophisticated approach, you can set up a separate visualization system that gives you an overview of all your ML models and datasets and provides an ongoing, updated view of metrics.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-c251cda4d25de23d6d1f48c9249a64aa5d12dd8a%2F2023109_course_module1_fin_images.066-min.png?alt=media)

## Summing up

Each ML monitoring architecture has its pros and cons. When choosing between them, consider existing tools, the scale of ML deployments, and available team resources for systems support. Be pragmatic: you can start with a simpler architecture and expand later.

For a deeper dive into the ML monitoring architectures with specific code examples, head to [Module 5](/ml-observability-course/module-5-ml-pipelines-validation-and-testing) and [Module 6](/ml-observability-course/module-6-deploying-an-ml-monitoring-dashboard).

## Enjoyed the content?

Star Evidently on GitHub to contribute back! This helps us create free, open-source tools and content for the community.

⭐️ [Star](https://github.com/evidentlyai/evidently) on GitHub!


# Module 2: ML monitoring metrics

This module covers different aspects of the production ML model performance. You will learn how to apply data quality, model quality, and data drift metrics for structured data.

This module will cover different aspects of the production ML model performance. We will explain some popular metrics and tests and how to apply them:

* what it means to have a “good” ML model;
* evaluating ML model quality;
* tracking data quality in production;
* data and prediction drift as proxy metrics.

This module includes both theoretical parts and code practice for each evaluation type. At the end of this module, you will understand the contents of ML observability: metrics and checks you can run and how to interpret them.


# 2.1. How to evaluate ML model quality

How to evaluate ML model quality directly and use early monitoring to detect potential ML model issues.

{% embed url="<https://www.youtube.com/watch?v=7Y819MAQTDg&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=7>" %}

**Video 1**. [How to evaluate ML model quality](https://www.youtube.com/watch?v=7Y819MAQTDg\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=7), by Emeli Dral

## Challenges of standard ML monitoring

When it comes to standard ML monitoring, we usually start by measuring ML model performance metrics:

* **Model quality and error metrics** show how the ML model performs in production. For example, you can track precision, recall, and log-loss for classification models or MAE for regression models.
* **Business or product metrics** help evaluate the ML model’s impact on business performance. You might want to track such metrics as purchases, clicks, views, etc.

**However, standard ML monitoring is not always enough**. Some challenges can complicate the ML performance assessment:

* **Feedback or ground truth is delayed**. When ground truth is not immediately available, calculating quality metrics can be technically impossible.
* **Past performance does not guarantee future results**, especially when the environment is unstable.
* **Many segments with different quality**. Aggregated metrics might not provide insights for diverse user/object groups. In this case, we need to monitor quality metrics for each segment separately.
* **The target function is volatile**. Volatile target function can lead to fluctuating performance metrics, making it difficult to differentiate between local quality drops and major performance issues.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-535ca4b5b38ecd8ac5f29640a4f43caa40b3268b%2F2023109_course_module2.005-min.png?alt=media)

## Early monitoring metrics

You can adopt **early monitoring** together with standard monitoring metrics to tackle these challenges.

Early monitoring focuses on metrics derived from consistently available data: input data and ML model output data. For example, you can track:

* **Data quality** to detect issues with data quality and integrity.
* **Data drift** to monitor changes in the input feature distributions.
* **Output drift** to observe shifts in model predictions.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a209f78239852f226627cf9370279c9c341a2bcb%2F2023109_course_module2.006-min.png?alt=media)

## Module 2 structure

This module includes both theoretical parts and code practice for each of the evaluation types. Here is the module structure:

**Model quality**

* Theory: ML model quality metrics for regression, classification, and ranking problems.
* Practice: building a sample report in Python showcasing quality metrics.

**Data quality**

* Theory: data quality metrics.
* Practice: creating a sample report in Python on data quality.

**Data and prediction drift**

* Theory: an overview of the data drift metrics.
* \[OPTIONAL] Theory: a deeper dive into data drift detection methods and strategies.
* Practice: building a sample report in Python to detect data and prediction drift for various data type.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-bfe31bae240ab317ad22c0c19a36152d7dd13195%2F2023109_course_module2.007-min.png?alt=media)

## Summing up

Tracking ML quality metrics in production is crucial to ensure that ML models perform reliably in real-world scenarios. However, standard ML performance metrics like model quality and error are not always enough.

Adopting early monitoring and measuring data quality, data drift, and prediction drift provides insights into potential issues when standard performance metrics cannot be calculated.

Through this module, learners will gain a theoretical understanding and hands-on experience in evaluating and interpreting model quality, data quality, and data drift metrics.


# 2.2. Overview of ML quality metrics. Classification, regression, ranking

Commonly used ML quality metrics for classification, regression, and ranking problems.

{% embed url="<https://www.youtube.com/watch?v=4_LOXDWxCbw&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=8>" %}

**Video 2**. [Overview of ML quality metrics](https://www.youtube.com/watch?v=4_LOXDWxCbw\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=8), by Emeli Dral

## ML model quality in production

ML model quality degrades over time. This happens because things change, and the model’s environment evolves.

You need **monitoring** to be able to maintain the ML model's relevance by detecting issues on time. You can also collect additional data and build visualizations for **debugging**.

But there is a caveat: to calculate classification, regression, and ranking quality metrics, **you need labels**. If you can, consider labeling at least part of the data to be able to compute them.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-691ddc77ed9315bf39258f603e6071992f94d74a%2F2023109_course_module2.009-min.png?alt=media)

## Classification quality metrics

A classification problem in ML is a task of assigning predefined categories or classes (labels) to new input data. Here are some commonly used metrics to measure the quality of the classification model:

* [**Accuracy**](https://www.evidentlyai.com/classification-metrics/accuracy-precision-recall) is the overall share of correct predictions. It is well-interpretable and arguably the most popular metric for classification problems. However, be cautious when using this metric with imbalanced datasets.
* [**Precision**](https://www.evidentlyai.com/classification-metrics/accuracy-precision-recall) measures correctness when predicting the target class.
* [**Recall**](https://www.evidentlyai.com/classification-metrics/accuracy-precision-recall) shows the ability to find all the objects of the target class. Precision and recall are usually used together. Both work well for unbalanced datasets.
* **F1-score** is the harmonic mean of precision and recall.
* [**ROC-AUC**](https://www.evidentlyai.com/classification-metrics/explain-roc-curve) works for probabilistic classification and evaluates the model's ability to rank correctly.
* **Logarithmic loss** demonstrates how close the prediction probability is to the actual value. It is a good metric for probabilistic problem statement.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-e4b8c20f72ba14fea988a69638dc2584835708a9%2F2023109_course_module2.012-min.png?alt=media)

Methods to help visualize and understand classification quality metrics include:

* [**Confusion matrix**](https://www.evidentlyai.com/classification-metrics/confusion-matrix) shows the number of correct predictions – true positives (TP) and true negatives (TN) – and the number of errors – false positives (FP) and false negatives (FN). You can calculate precision, recall, and F1-score based on these values.
* **Precision-recall table** helps calculate metrics like precision, recall, and F1-score for different thresholds in probabilistic classification.
* **Class separation quality** helps visualize correct and incorrect predictions for each class.
* **Error analysis**. You can also map predicted probabilities or model errors alongside feature values and explore if a specific type of misclassification is connected to the particular feature values.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-b3757cd061348a0b0cc164250526a90d1be0432c%2F2023109_course_module2.016-min.png?alt=media)

{% hint style="info" %}
**Further reading:** [What is your model hiding? A tutorial on evaluating ML models](https://www.evidentlyai.com/blog/tutorial-2-model-evaluation-hr-attrition).
{% endhint %}

## Regression quality metrics

Regression models provide numerical output which is compared against actual values to estimate ML model quality. Some standard regression quality metrics include:

* **Mean Error (ME)** is an average of all errors. It is easy to calculate, but remember that positive and negative errors can overcompensate each other.
* **Mean Absolute Error (MAE)** is an average of all absolute errors.
* **Root Mean Squared Error (RMSE)** is a square root of the mean of squared errors. It penalizes larger errors.
* **Mean Absolute Percentage Error (MAPE)** averages all absolute errors in %. Works well for datasets with objects of different scale (i.e., tens, thousands, or millions).
* **Symmetric MAPE** provides different penalty for over- or underestimation.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-ccfb475ec02d59aeebf5dfd1b98c9e8a561140e2%2F2023109_course_module2.020-min.png?alt=media)

Some of the methods to analyze and visualize regression model quality are:

* **Predicted vs. Actual** value plots and Error over time plots help derive patterns in model predictions and behavior (e.g., Does the model tend to have bigger errors during weekends or hours of peak demand?).
* **Error analysis**. It is often important to distinguish between **underestimation** and **overestimation** during error analysis. Since errors might have different business costs, this can help optimize model performance for business metrics based on the use case.

You can also map extreme errors alongside feature values and explore if a specific type of error is connected to the particular feature values.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-abe416f5a8c2b46a6215bf87512881531b6f4136%2F2023109_course_module2.025-min.png?alt=media)

## Ranking quality metrics

Ranking focuses on the relative order of items rather than their absolute values. Popular examples of ranking problems are search engines and recommender systems.

We need to estimate the order of objects to measure quality in ranking tasks. Some commonly used ranking quality metrics are:

* **Cumulative gain** helps estimate the cumulative value of recommendations and does not take into account the position of a result in the list.
* **Discounted Cumulative Gain (DCG)** gives a penalty when a relevant result is further in the list.
* **Normalized DCG (NDCG)** normalizes the evaluation irrespective of the list length.
* **Precision @k** is a share of the relevant objects in top-K results.
* **Recall @k** is a coverage of all relevant objects in top-K results.
* **Lift @k** reflects an improvement over random ranking.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a1a18b7e7060f58e3dfff22d8e6f9e2ce22f1b96%2F2023109_course_module2.028-min.png?alt=media)

If you work on a recommender system, you might want to consider additional – “beyond accuracy” – metrics that reflect RecSys behavior. Some examples are:

* Serendipity
* Novelty
* Diversity
* Coverage
* Popularity bias

You can also use other custom metrics based on your problem statement and business context, for example, by weighting the metrics by specific segments.

## Considerations for production ML monitoring

When you define the model quality metrics to monitor the ML model performance in production, there are some important considerations to keep in mind:

**Pick the right metrics that align with your use case and business goals**:

* **The usuals apply**. E.g., reuse metrics from the model development phase and do not use accuracy for a problem with highly imbalanced classes.
* **Consider a proxy business metric to evaluate impact**. E.g., consider tracking an estimated loss/gain based on known error costs, the share of predictions with an error larger than X, etc.
* **Not all evaluation metrics are useful for dynamic production monitoring**. E.g., ROC AUC reflects quality across all thresholds, but a production model has a specific one.
* **Consider custom metrics or heuristics**. E.g., the average position of the first relevant object in the recommendation block.

**It’s not just a choice of metric**. There are other parameters you might need to define:

* **Aggregation window**. It is crucial to calculate metrics in the right windows. Depending on the use case, you might want to monitor precision, for example, every minute, hourly, daily, or over a sliding 7-day window as a key performance indicator.
* **Segments**. You can track model quality separately for different locations, devices, customer subscription types, etc.

## Summing up

We discussed the importance of monitoring ML model performance in production and introduced commonly used quality metrics for classification, regression, and ranking problems.

Further reading: [What is your model hiding? A tutorial on evaluating ML models](https://www.evidentlyai.com/blog/tutorial-2-model-evaluation-hr-attrition).

In the next part of this module, we will dive into practice and build a model quality report using the open-source [Evidently](https://github.com/evidentlyai/evidently) Python library.


# 2.3. Evaluating ML model quality \[CODE PRACTICE]

A code example walkthrough of ML model quality evaluation using Python and the open-source Evidently library.

{% embed url="<https://www.youtube.com/watch?v=QWLw_lJ29k0&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=9>" %}

**Video 3**. [Evaluating ML model quality \[CODE PRACTICE\]](https://www.youtube.com/watch?v=QWLw_lJ29k0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=9), by Emeli Dral

In this video, we walk you through the code example of ML model quality evaluation using Python and the open-source [Evidently](https://github.com/evidentlyai/evidently) library.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module2/ml_model_quality.ipynb) to follow along.

**Outline**:\
[00:00](https://www.youtube.com/watch?v=QWLw_lJ29k0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=9\&t=0s) Create a working environment and import libraries\
[02:45](https://www.youtube.com/watch?v=QWLw_lJ29k0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=9\&t=165s) Prepare datasets for classification and regression models\
[08:25](https://www.youtube.com/watch?v=QWLw_lJ29k0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=9\&t=505s) Build and customize classification quality report\
[14:50](https://www.youtube.com/watch?v=QWLw_lJ29k0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=9\&t=890s) Save and share the report\
[16:05](https://www.youtube.com/watch?v=QWLw_lJ29k0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=9\&t=965s) Display the report in JSON format and as a Python dictionary\
[18:15](https://www.youtube.com/watch?v=QWLw_lJ29k0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=9\&t=1095s) Build and customize regression quality report

That’s it! We built an ML model quality report for classification and regression problems and learned how to display it in HTML and JSON formats and as a Python dictionary.


# 2.4. Data quality in machine learning

Types of production data quality issues, how to evaluate data quality, and interpret data quality metrics.

{% embed url="<https://www.youtube.com/watch?v=IRbmQGqzVZo&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=10>" %}

**Video 4**. [Data quality in machine learning](https://www.youtube.com/watch?v=IRbmQGqzVZo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=10), by Emeli Dral

## What can go wrong with the input data?

If you have a complex ML system, there are many things that can go wrong with the data. The golden rule is: garbage in, garbage out. We need to make sure that the data we feed our model with is fine.

Some common data processing issues are:

* **Wrong source**. E.g., a pipeline points to an older version of the table.
* **Lost access**. E.g., permissions are not updated.
* **Bad SQL. Or not SQL**. E.g., a query breaks when a user comes from a different time zone and makes an action “tomorrow."
* **Infrastructure update**. E.g., change in computation based on a dependent library.
* **Broken feature code**. E.g., feature computation breaks at a corner case like a 100% discount.

Issues can also arise if the data schema changes or data is lost at the source (e.g., broken in-app logging or frozen sensor values). If you have several models interacting with each other, broken upstream models can affect downstream models.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-40327510a88612e769f9005652a4081e6e608664%2F2023109_course_module2.041-min.png?alt=media)

## Data quality metrics and analysis

**Data profiling** is a good starting point for monitoring data quality metrics. Based on the data type, you can come up with basic descriptive statistics for your dataset. For example, for numerical features, you can calculate:

* Min and Max values
* Quantiles
* Unique values
* Most common values
* Share of missing values, etc.

Then, you can visualize and compare statistics and data distributions of the current data batch and reference data to ensure data stability.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-d2953dd9ba2cae5a65f7a51f4326f32cabe310ff%2F2023109_course_module2.047-min.png?alt=media)

When it comes to monitoring data quality, you must define the conditions for alerting.

**If you do not have reference data, you can set up thresholds manually based on domain knowledge**. “General ML data quality” can include such characteristics as:

* no/low share of missing values
* no duplicate columns/rows
* no constant (or almost constant!) features
* no highly correlated features
* no target leaks (high correlation between feature and target)
* no range violations (based on the feature context, e.g., negative age or sales).

Since setting up these conditions manually can be tedious, it often helps to have a reference dataset.

**If you have reference data, you can compare it with the current data and autogenerate test conditions based on the reference**. For example, based on the training or past batch, you can monitor for:

* expected data schema and column types
* expected data completeness (e.g., 90% non-empty)
* expected batch size (e.g., number of rows)
* expected patterns for specific columns, such as:
  * non-unique (features) or unique (IDs)
  * specific data distribution types (e.g., normality)
  * expected ranges based on observed values
  * descriptive statistics: averages, median, quantiles, min-max (point estimation or statistical tests with a confidence interval).

## Summing up

Monitoring data quality is critical to ensuring that ML models function reliably in production. Depending on the availability of reference data, you can manually set up thresholds based on domain knowledge or automatically generate test conditions based on the reference.

Up next: hands-on practice on how to evaluate and test data quality using Python and [Evidently](https://github.com/evidentlyai/evidently) library.


# 2.5. Data quality in ML \[CODE PRACTICE]

A code example walkthrough of data quality evaluation using Evidently Reports and Test Suites.

{% embed url="<https://www.youtube.com/watch?v=_HKGrW2mVdo&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=11>" %}

**Video 5**. [Data quality in ML \[CODE PRACTICE\]](https://www.youtube.com/watch?v=_HKGrW2mVdo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=11), by Emeli Dral

In this video, we walk you through the code example of data quality evaluation using [Evidently](https://github.com/evidentlyai/evidently) Reports and Test Suites.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module2/data_quality.ipynb) to follow along.

Here is a quick refresher on the Evidently components we will use:

* **Reports** compute and visualize 100+ metrics in data quality, drift, and model performance. You can use in-built report presets to make visuals appear with just a couple of lines of code.
* **Test Suites** perform structured data and ML model quality checks. They verify conditions and show which of them pass or fail. You can start with default test conditions or design your testing framework.

**Outline**:\
[00:00](https://www.youtube.com/watch?v=_HKGrW2mVdo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=11\&t=0s) Create a working environment and import libraries\
[01:30](https://www.youtube.com/watch?v=_HKGrW2mVdo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=11\&t=90s) Prepare reference and current dataset\
[05:20](https://www.youtube.com/watch?v=_HKGrW2mVdo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=11\&t=320s) Run data quality Test Suite and visualize the results\
[09:30](https://www.youtube.com/watch?v=_HKGrW2mVdo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=11\&t=570s) Customize the Test Suite by specifying individual tests and test conditions\
[13:20](https://www.youtube.com/watch?v=_HKGrW2mVdo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=11\&t=800s) Build and customize data quality Report

That’s it! We evaluated data quality using Evidently Reports and Test Suites and demonstrated how to add custom metrics, tests, and test conditions to the analysis.


# 2.6. Data and prediction drift in ML

What data and prediction drift is, and how to detect distribution drift using statistical methods and rule-based checks.

{% embed url="<https://www.youtube.com/watch?v=bMYcB_5gP4I&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=12>" %}

**Video 6**. [Data and prediction drift in ML](https://www.youtube.com/watch?v=bMYcB_5gP4I\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=12), by Emeli Dral

## What is data drift, and why evaluate it?

When ground truth is unavailable or delayed, we cannot calculate ML model quality metrics directly. Instead, we can use proxy metrics like feature and prediction drift.

**Prediction drift** shows changes in the distribution of **model outputs** over time. Without target values, this is the best proxy of the model behavior. Detected changes in the model outputs may be an early signal of changes in the model environment, data quality bugs, pipeline errors, etc.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-e097574f77475cf89f682c071ed585391740734e%2F2023109_course_module2.058-min.png?alt=media)

**Feature drift** demonstrates changes in the distribution of **input features** over time. When we train the model, we assume that if the input data remains reasonably similar, we can expect similar model quality. Thus, data distribution drift can be an early warning about model quality decay, important changes in the model environment or user behavior, unannounced changes to the modeled process, etc.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-11b72697b5af9f1d4940867d8eca1c7a20fdefd5%2F2023109_course_module2.060-min.png?alt=media)

Prediction and feature drift can serve as early warning signs for model quality issues. They can also help pinpoint a root cause when the model decay is already observed.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-adf6edb2608018da689d995cfb0005d0da4c9f58%2F2023109_course_module2.065-min.png?alt=media)

Some key considerations about data drift to keep in mind:

* **Prediction drift is usually more important than feature drift**. If you monitor one thing, look at the outputs.
* **Data drift in ML is a heuristic**. There is no “objective” drift; it varies based on the specific use case and data.
* **Not all distribution drift leads to model performance decay**. Consider the use case, the meaning of specific features, their importance, etc.
* **You don’t always need to monitor data drift**. It is useful for business-critical models with delayed feedback. But often you can wait.
* **Data drift helps with debugging**. Even if you do not alert on feature drift, it might help troubleshoot the decay.
* **Drift detection might be valuable even if you have the labels**. Feature drift might appear before you observe the model quality drop.

{% hint style="info" %}
**Further reading:** [How to break a model in 20 days. A tutorial on production model analytics](https://www.evidentlyai.com/blog/tutorial-1-model-analytics-in-production).
{% endhint %}

## How to detect data drift?

To detect distribution drift, you need to pick:

* **Drift detection method**: statistical tests, distance metrics, rules, etc.
* **Drift detection threshold**: e.g., confidence levels for statistical tests or numeric threshold for distance metrics.
* **Reference dataset**: what an exemplary distribution is.
* **Alert conditions**: e.g., based on feature importance and the share of the drifting features.

## Data drift detection methods

There are three commonly used approaches to drift detection:

* **Statistical tests**, e.g., Kolmogorov-Smirnov or Chi-squared test. You can use parametric or non-parametric tests to compare distributions. Generally, parametric tests are more sensitive. Using statistical tests for drift detection is best for smaller datasets and samples. The resulting drift “score” is measured by p-value (a “confidence” of drift detection).
* **Distance-based metrics**, e.g., Wasserstein distance or Jensen Shannon Divergence. This group of metrics works well for larger datasets. The drift “score” is measured as distance, divergence, or level of similarity.
* **Rule-based checks** are custom rules for detecting drift based on heuristics and domain knowledge. These are great when you expect specific changes, e.g., new categories added to the dataset.

Here is how the defaults are implemented in the Evidently open-source library.

**For small datasets (<=1000)**, you can use Kolmogorov-Smirnov test for numerical features, Chi-squared test for categorical features, and proportion difference test for independent samples based on Z-score for binary categorical features.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-526ee287826350bf287e18bf427c478fa56c766e%2F2023109_course_module2.070-min.png?alt=media)

**For large datasets (>1000)**, you might use Wasserstein Distance for numerical features and Jensen-Shannon divergence for categorical features.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-866f4f68b956a95cd00be5c1e215fe551c47df1a%2F2023109_course_module2.071-min.png?alt=media)

## Univariate vs. multivariate drift

The **univariate drift** detection approach looks at drift in each feature individually. It returns drift/no drift for each feature and can be easily interpretable.

The **multivariate drift** detection approach looks at the complete dataset (e.g., using PCA and certain methods like domain classifier). It returns drift/no drift for the dataset and may be useful for systems with many features.

You can still use the univariate approach to detect drift in a dataset by:

* Tracking the share (%) of drifting features to get a dataset drift decision.
* Tracking distribution drift only in the top model features.
* Combining both solutions.

## Tips for calculating drift

Here are some tips to keep in mind when calculating data drift:

* **Data quality is a must**. Calculate data quality metrics first and then monitor for drift. Otherwise, you might detect “data drift” that is caused by data quality issues.
* **Mind the feature set**. The approach to drift analysis varies based on the type and importance of features.
* **Mind the segments**. Consider segment-based drift monitoring when you have clearly defined segments in your data. For example, in manufacturing, you might have different suppliers of raw materials and need to monitor distribution drift separately for each of them.

## Summing up

We discussed the key concepts of data drift and how to measure it. When calculating data drift, consider drift detection method and thresholds, properties of reference data, and alert conditions.

Further reading: [How to break a model in 20 days. A tutorial on production model analytics](https://www.evidentlyai.com/blog/tutorial-1-model-analytics-in-production).

Up next: deep dive into data drift detection \[OPTIONAL] and practice on how to detect data drift using Python and [Evidently](https://github.com/evidentlyai/evidently) library.


# 2.7. Deep dive into data drift detection \[OPTIONAL]

A deep dive into data drift detection methods, how to choose the right approach for your use case, and what to do when the drift is detected.

{% embed url="<https://www.youtube.com/watch?v=N47SHSP6RuY&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=13>" %}

**Video 7**. [Data and prediction drift in ML](https://www.youtube.com/watch?v=N47SHSP6RuY\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=13), by Emeli Dral

Welcome to the deep dive into data drift detection! We will cover the following topics:

**Data drift detection methods:**

* More on methods to detect data drift
* Strategies for choosing drift detection approach

**Special cases:**

* Detecting drift for large datasets
* Detecting drift for real-time models
* Using drift as a retraining trigger

**Useful tips:**

* How to interpret prediction and data drift together?
* What to do after drift is detected

## Drift detection methods

Let’s have a closer look at the commonly used approaches to drift detection.

**Parametric statistical tests** There are both **one-sample** and **two-sample** parametric tests.:

* **If you only have current data** and no reference data is available, you can use **Z-test and T-test for mean** (m = m0) or **one-proportion Z-test** (p = p0) to detect data drift. These methods can work if you have interpretable features – e.g., salary or age – as you need to develop the hypotheses on distribution values.
* **If reference data is available**, you can use two-sample parametric tests to compare distributions: for example, **two-proportions Z-test** or **two-sample Z-test and T-test for means** (for normally distributed samples).

Some considerations to keep in mind when using parametric tests:

* **They require different tests for different features.** For example, some tests assume that your data are normally distributed.
* **They are more sensitive to drift than non-parametric tests.** If you work on a problem where you have a small dataset and want to react to even minor deviations – it makes sense to use parametric tests.
* **Hard to fine-tune if you have a lot of features.** If you have many features with different feature types, you need to invest a lot of time in choosing the right test for each feature.

It makes sense to use parametric tests if you have a small number of interpretable features and work on critical use cases (e.g., in healthcare).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-129ef1061c21882c3fcda46f5a1f70d2512e7472%2F2023109_course_module2.081-min.png?alt=media)

**Non-parametric statistical tests** Non-parametric tests are less demanding to the properties of data samples and thus are widely used. Examples include **Kolmogorov-Smirnov** test, K-sample **Anderson-Darling** test, **Pearson’s chi-squared** test, **Fisher’s/Barnard’s** exact test for small samples, etc.

When using non-parametric tests, consider the following:

* **Feature type.** You can use heuristics to choose suitable tests based on the feature type, e.g., numerical, categorical, or binary.
* **Sensitivity.** Non-parametric tests are less sensitive to drift than parametric tests.
* **Data volumes.** It makes sense to use non-parametric tests for low-volume datasets or samples (e.g., less than 1000 objects).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-13c86cb926cc477800ad1b73dfa14df8d12a6cb4%2F2023109_course_module2.082-min.png?alt=media)

**Distance-based approaches** Distance-based methods measure how far two distributions are from each other and thus are easy to interpret. For example, you can calculate **Wasserstein distance**, **Jensen-Shannon divergence**, or **Population Stability Index** (PSI).

Some considerations to keep in mind when using distance-based methods to detect data drift:

* **Variety of metrics is available.** Roughly any metric that shows difference/similarity between distributions can be used as a drift detection method.
* **High interpretability compared to statistical tests.** Often, it makes more sense to pick an interpretable metric rather than a statistical test.
* **Data volume.** It makes sense to use distance-based methods for larger datasets (e.g., > 1000 objects).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-eddce83b31f98a1cd9a7793fb013b162164e1eca%2F2023109_course_module2.083-min.png?alt=media)

**Domain classification** This approach uses binary classifiers to distinguish between reference and current data. It can be used to detect data drift in different data types, including embeddings, unstructured data (such as texts), and multimodal data.

{% hint style="info" %}
**Further reading:** [Which test is the best? We compared 5 methods to detect data drift on large datasets](https://www.evidentlyai.com/blog/data-drift-detection-large-datasets).
{% endhint %}

## How to choose a drift detection approach?

**Data drift detection is a heuristic.** There is no strict law. This is why it is important that you consider why you want to detect drift, and what method might make sense for you.

Consider your problem statement and dataset properties. For example:

* If your use case is sensitive, you might want to use parametric tests.
* If interpretability is important, consider distance-based methods.
* Domain classification can be a good choice if you work with various data types – text, videos, tabular data.

To choose the right drift detection approach for your particular problem statement, you can consider two options:

**1. Go with defaults.**

In this scenario, you pick some reasonable defaults to start and adjust the sensitivity as you proceed with monitoring.

* Start with basic assumptions. Do you want to detect drift for the whole dataset or only consider drift in important features?
* Pick reasonable metrics and thresholds. For example, for numerical features, you can pick Wasserstein Distance at 0.1 threshold.
* Start monitoring.
* Visualize results.
* Adjust based on false alarms, sensitivity, and drift interpretations.

**2. Experiment.**

In this scenario, you use historical data to tweak detection parameters using past known drifts. Here is an example of an experiment:

* Take data for a stable period
* Take data with known drift or simulate drift using synthetic data.
* Apply different drift detection approaches. Experiment with tests, thresholds, window size and/or bucketing parameters.
* Choose the optimal approach that detects known drifts and minimizes false alarms.

## Special cases

There are some special cases to keep in mind when detecting data drift:

**Large datasets.** Statistics was made to work with samples. Having many objects and/or features in a dataset can lead to some tests being “too sensitive” or taking too long to compute. If this is the case, you can use **sampling** to pick representative observations and apply tests on top of them. Alternatively, you can try **bucketing** to aggregate observations and reduce the amount of data. For example, you can detect drift on top of hourly data instead of minutely.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-5dbbc3d1d5c0dd11c4b1392a18ff3bf4d6f82808%2F2023109_course_module2.089-min.png?alt=media)

**Non-batch models.** While some metrics can be calculated in real-time, we need to generate a batch of data to detect data drift.

The solution is to use **window functions** to perform tests on continuous data streams. You can pick a window function (i.e., moving windows with/without moving reference), choose the window and step size to create batches for comparison, and “compare” the windows.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-00bf5bed6697cebd13630993469b8710c1073e35%2F2023109_course_module2.090-min.png?alt=media)

**Feature drift as a retraining trigger.** There are both pros and cons of using drift detection as the retraining trigger.

Generally, we do not recommend retraining a model every time the drift is detected because:

* **Data might be low-quality.** Retraining the model on corrupted data will be useless if data drift occurs due to data processing issues.
* **Data might be insufficient.** Sometimes, we just don’t have enough data for new model training.
* **Data might be non-representative.** Look out for unstable periods, e.g., pandemic, seasonal spikes, etc.

Instead, try to understand data drift first:

* **Data drift as an investigation trigger.** Try to figure out the root cause of the detected drift.
* **Data drift as a labeling trigger.** You can use a data drift signal to start the labeling process to be able to compute the actual model quality metrics.

If you use data drift as a retraining trigger, it is critical to implement a solid evaluation process before roll-out to make sure the new model performs well.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-0bc8a93ebbad249c625d1163b36a5f8966a2e45f%2F2023109_course_module2.091-min.png?alt=media)

## How to interpret data and prediction drift together?

It often makes sense to monitor both prediction drift (change in the model outputs) and data drift (change in the model features).

However, data and prediction drift do not necessarily mean that something is wrong. Let’s look at two examples of data and prediction drift detected together or independently.

**Scenario 1. Data drift: detected. Prediction drift: not detected.** There are both positive and negative ways to interpret it.

**Positive interpretation:**

* Important features did not change.
* Model is robust enough to survive drift.
* No need to intervene.

**Negative interpretation:**

* Important features changed.
* Model should have reacted but did not. It does not extrapolate well.
* We need to intervene.

**Scenario 2. Data drift: detected. Prediction drift: detected.** Again, there are positive and negative ways of interpreting it.

**Positive interpretation:**

* Important features changed.
* Model reacts and extrapolates well (e.g., prices lower -> higher sales)
* No need to intervene.

**Negative interpretation:**

* Important features changed.
* Model behavior is unreasonable.
* We need to intervene.

## What to do if drift is detected?

Here are some possible steps to take if the drift is detected:

**1. Check the data quality.** Make sure the drift is “real” and try to interpret where the drift is coming from. Data entry errors, stale features, and lost data are data quality issues disguised as data drift. If this is the case, fix the data first.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-80ec18e603a7e07d8b4f49b488ac319b83c1f39d%2F2023109_course_module2.098-min.png?alt=media)

**2. Investigate the drift.** Analyze which features have changed and how much. To understand the shift, you can:

* Visualize distributions
* Analyze correlation changes
* Check descriptive stats
* Evaluate segments
* Seek real-world explanations (e.g., a new marketing campaign)
* Team up with domain experts

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a39fe07d2f133d87dcfa93b76819c6b0e2858513%2F2023109_course_module2.100-min.png?alt=media)

**3. Doing nothing is also an option.** You might treat the drift as a false alarm, be satisfied with how the model handles drift, or simply decide to wait.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-af8c50cfb691240faca8038aab7afcdb72bc0d46%2F2023109_course_module2.101-min.png?alt=media)

**4. Actively reacting to drift.** However, often, we need to react when the drift is detected:

* **Retrain the model.** Get new labels and actual values and re-fit the same model on the latest data.
* **Rebuild the model.** If the change is significant, you might need to rebuild the training pipeline and test new model architectures.
* **Tune the model.** For example, you can change a threshold for drift detection.
* **Use a fallback strategy.** Decide without ML: switch to manual processing, heuristics, or non-ML models.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-065d6d89013d2340549a639b85a25e4f7e06aab1%2F2023109_course_module2.102-min.png?alt=media)

{% hint style="info" %}
**Further reading:** ["My data drifted. What's next?" How to handle ML model drift in production.](https://www.evidentlyai.com/blog/ml-monitoring-data-drift-how-to-handle).
{% endhint %}

## Summing up

We discussed different drift detection methods and how to choose the optimal approach for your dataset and problem statement. We covered special cases like handling large datasets, calculating drift for real-time models, and using drift as a retraining trigger. We also learned how to interpret data and prediction drift and what to do if drift is detected.

Further reading:

* [Which test is the best? We compared 5 methods to detect data drift on large datasets](https://www.evidentlyai.com/blog/data-drift-detection-large-datasets)
* ["My data drifted. What's next?" How to handle ML model drift in production.](https://www.evidentlyai.com/blog/ml-monitoring-data-drift-how-to-handle)

Up next: code practice on how to detect data drift using the open-source [Evidently](https://github.com/evidentlyai/evidently) Python library.


# 2.8. Data and prediction drift in ML \[CODE PRACTICE]

A code example walkthrough of detecting data drift and creating a custom method for drift detection using Evidently.

{% embed url="<https://www.youtube.com/watch?v=oO1K4CaWxt0&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=14>" %}

**Video 8**. [Data and prediction drift in ML \[CODE PRACTICE\]](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14), by Emeli Dral

In this video, we walk you through the code example of detecting data drift and creating a custom method for drift detection using the open-source [Evidently](https://github.com/evidentlyai/evidently) Python library.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module2/data_drift_deep_dive.ipynb) to follow along.

**Outline**:\
[00:00](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14\&t=0s) Create a working environment and import libraries\
[01:33](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14\&t=93s) Overview of the data drift options\
[04:25](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14\&t=265s) Evaluating share of drifted features\
[06:40](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14\&t=400s) Detecting column drift\
[11:47](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14\&t=707s) Set different drift detection method per feature type\
[12:57](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14\&t=777s) Set individual different drift detection methods per feature\
[15:34](https://www.youtube.com/watch?v=oO1K4CaWxt0\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=14\&t=934s) Custom drift detection method

## Enjoyed the content?

Star Evidently on GitHub to contribute back! This helps us create free, open-source tools and content for the community.

⭐️ [Star](https://github.com/evidentlyai/evidently) on GitHub!


# Module 3: ML monitoring for unstructured data

This module covers evaluating and monitoring the production performance for models that use unstructured data, including NLP, LLMs, and embeddings.

This module covers evaluating and monitoring the production performance for models that use unstructured data, including LLM-based systems.

We will cover:

* Why monitoring unstructured data is difficult;
* How to measure text data quality;
* What are text descriptors and how to use them;
* How to deal with embeddings;
* How to deal with multimodal data.

This module includes both a **theoretical part and code practice**. At the end of this module, you will understand the possible approaches to monitoring ML models that work with texts and other unstructured data.


# 3.1. Introduction to NLP and LLM monitoring

How to evaluate model quality for NLP and LLMs and monitor text data without labels using raw data, embeddings, and descriptors.

{% embed url="<https://www.youtube.com/watch?v=tYA0h3mPeZE&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=15>" %}

**Video 1**. [Introduction to NLP and LLM monitoring](https://www.youtube.com/watch?v=tYA0h3mPeZE\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=15), by Emeli Dral

## Use cases for NLP and LLM models

NLP and LLM models are widely used for both predictive and generative tasks.

**Predictive applications** include:

* **Classification tasks**, e.g., spam detection, classification of support tickets, evaluating text sentiment, etc.
* **Search (ranking) tasks** are common in e-commerce website search, content recommendations, etc.
* **Information extraction**, like extracting structured information (names, dates, etc.) from text, etc.

**Generative applications** cover such tasks as translation, summarization, and text generation, e.g., conversational interfaces, article generation, code generation, etc.

Modern ML systems are often a combination of those: e.g., a support chatbot includes a classifier to detect intent and a generative model to produce a text response.

## What can go wrong with NLP and LLMs?

All production ML models need monitoring. NLP and LLM models are no exception. To build our monitoring strategy, we need to understand what can go wrong with the models that run on the unstructured data.

There are standard issues that apply to both unstructured and tabular data:

* **Technical errors** such as wrong encoding or data processing bugs.
* **Data shifts** like new topics or unexpected usage scenarios.

However, there are also issues specific to working with unstructured data. For example:

* **Model attacks,** e.g., prompt injection and adversarial usage.
* **Model behavior shift.** For example, if you rely on third-party models, you can face changes in their properties due to retraining.
* **Hallucinations** arise when the model starts to generate factually incorrect or unrelated answers.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-296e7fbae325c192cc80076abc0d0affdeef36f0%2F2023109_course_module3.004-min.png?alt=media)

## ML monitoring metrics for NLP and LLMs

So, what can you monitor to detect the discussed issues? There are two groups of signals to keep tabs on:

**1. Direct signals.** This group includes standard model quality metrics such as accuracy, precision, and recall – and requires labeled data.

When dealing with **predictive applications** – e.g., classification or ranking tasks – we typically have a "right" answer. In this case, we can label the data and then compare labels and the model’s outputs to calculate direct quality metrics.

However, for **generative applications** (e.g., translation, text summarization), there is often no single correct answer and many possible "good" responses. If this is the case, relying on labels and standard quality metrics is not always possible. Instead, you can use feedback, human labeling, LLM-based grading, and response validation as proxy signals.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-697418e346fa25c9935f4474d5874a2c40a6802e%2F2023109_course_module3.006-min.png?alt=media)

**2. Proxy metrics** help evaluate the model properties. You can look out for data quality, data and prediction drift, or user feedback.

Here are three proxy strategies for unstructured data you can use *(disclaimer: we use text as the main type of unstructured data)*:

* **Monitoring raw text data.** Using analytical methods that process raw texts allows you to catch an interpretable signal and use it for hypothesis formulation and debugging.
* **Monitoring text descriptors.** You can extract signals from raw data and compute metrics based on these structured signals.
* **Monitoring embeddings.** When raw data is not available, the embeddings can be used. You can monitor for changes in the distributions of input and output embeddings.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-6d8e799cb09c58295dd77f0d46b33288401b0f32%2F2023109_course_module3.007-min.png?alt=media)

## Summing up

Monitoring NLP and LLM models differs from monitoring models built on structured data. Both predictive and generative applications have their own monitoring needs, and understanding these differences is crucial for effective model observability. You can use both direct quality metrics and proxy signals to keep tabs on production NLP- and LLM-based systems.

Up next: delving deeper into detecting raw text data drift.


# 3.2. Monitoring data drift on raw text data

How to detect and evaluate raw text data drift using domain classifier and topic modeling.

{% embed url="<https://www.youtube.com/watch?v=wHyXSyVg5Ag&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=16>" %}

**Video 2**. [Monitoring data drift on raw text data](https://www.youtube.com/watch?v=wHyXSyVg5Ag\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=16), by Emeli Dral

## Challenges of monitoring raw text data

Handling raw text data is more complex than dealing with structured tabular data. With structured data, you can usually define “good” or “expected” data, e.g., particular feature distributions or statistical values can signal the data quality. For unstructured data, there is no straightforward way to define data quality or extract a signal for raw text data.

When it comes to data drift detection, you can use two strategies that rely on raw text data: **domain classifier** and **topic modeling**.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-1d4436831c537f2d64ef8ec474b42938c8ad84fb%2F2023109_course_module3.011-min.png?alt=media)

## Domain classifier

**Domain classifier** method or **model-based drift detection** uses a classifier model to compare distributions of reference and current datasets by training a model that predicts to which dataset a specific text belongs. If the model can confidently identify which text samples belong to the current or reference dataset, the two datasets are probably sufficiently different.

{% hint style="info" %}
**Further reading:** this approach is described in more detail in the paper ["Failing loudly: An Empirical Study of Methods for Detecting Dataset Shift"](https://arxiv.org/abs/1810.11953).
{% endhint %}

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-acd6509c082725800fef11571b85c9bc061f06d9%2F2023109_course_module3.013-min.png?alt=media)

You can directly use the ROC AUC of the binary classifier as the “drift score” when you deal with **large datasets**. If you work with **smaller datasets** (< 1000), you can compare the model ROC AUC against a random classifier.

The benefit of using model-based drift detection on raw data is its **interpretability**. In this case, you can identify top words and text examples that were easy to classify to explain the drift and debug the model.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-0330a5aa0312794645df8ea0e4010b2858276387%2F2023109_course_module3.014-min.png?alt=media)

{% hint style="info" %}
**Further reading:** [Monitoring NLP models in production: a tutorial on detecting drift in text data](https://www.evidentlyai.com/blog/tutorial-detecting-drift-in-text-data).
{% endhint %}

## Topic modeling

Another strategy for evaluating raw data quality is **topic modeling**. The goal here is to categorize text into interpretable topic clusters, so instead of a binary classification model, we use a **clustering model**.

How it works:

* Apply the clustering model to new batches of data.
* Monitor the size and share of different topics over time.
* Changes in topics can indicate data drift.

Using this method can be challenging due to difficulties in building a good clustering model:

* There is no ideal structure in clustering.
* Typically, it requires manual tuning to build accurate and interpretable clusters.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-5aa46769ddb5b707f859d2d25d22889138a11d78%2F2023109_course_module3.018-min.png?alt=media)

## Summing up

Defining data quality and tracking data drift for text data can be challenging. However, we can extract interpretable signals from text data to detect drift. You can use such methods as domain classifier and topic modeling to monitor for drift and evaluate the quality of raw text data.

Further reading:

* [Failing loudly: An Empirical Study of Methods for Detecting Dataset Shift](https://arxiv.org/abs/1810.11953)
* [Monitoring NLP models in production: a tutorial on detecting drift in text data](https://www.evidentlyai.com/blog/tutorial-detecting-drift-in-text-data)

Up next: an exploration of alternative text drift detection methods that use descriptors.


# 3.3. Monitoring text data quality and data drift with descriptors

What text descriptors are and how to use them to monitor text data quality and data drift.

{% embed url="<https://www.youtube.com/watch?v=UwWGxyCHQSw&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=17>" %}

**Video 3**. [Monitoring text data quality and data drift with descriptors](https://www.youtube.com/watch?v=UwWGxyCHQSw\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=17), by Emeli Dral

## What is a text descriptor?

A **text descriptor** is any feature you can derive or calculate from raw text. Text descriptor transforms unstructured data (e.g., text) into structured data (e.g., numeric or categorical descriptors).

Depending on your goal and problem statement, there are various groups of descriptors you can calculate:

* **Text data quality.** Example descriptors are text length, the share of out-of-vocabulary words, the share of non-letter characters, regular expressions match, etc.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-29b8647f22dd672b6e2e6ab67b9764200edce2a8%2F2023109_course_module3.024-min.png?alt=media)

* **Text contents.** For instance, you can monitor the presence of trigger words like mentions of specific brands or competitors or text sentiment. Collaborating with product managers or business teams is recommended to determine the meaningful text descriptors for your problem statement and domain area.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-41e30a5644a7ecd9a985cfc688e068dfabd46091%2F2023109_course_module3.025-min.png?alt=media)

You can also use **LLM-based grading** that replicates the evaluation that human assessors can do. In this scenario, you automate evaluations by using an LLM to “judge” the generated outputs.

For example, it can assign a score, check whether the text answers a specific question, define the tone of the output, etc. This method has its caveats: it is not always reliable, can be costly, and requires another set of prompts and outputs to monitor.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-d1956884dee50d5251fbadc85f292b10de05d9de%2F2023109_course_module3.026-min.png?alt=media)

{% hint style="info" %}
**Further reading:** [Monitoring unstructured data for LLM and NLP with text descriptors](https://www.evidentlyai.com/blog/unstructured-data-monitoring).
{% endhint %}

## Monitoring text descriptors

Once you have a structured representation of the unstructured data – in our case, text descriptors – you can use monitoring techniques applied to structured data. For example, you can:

* Measure **descriptor distribution drift**.
* Run **rule-based checks**, such as min-max ranges for text length, the expected share of non-letter symbols, or responses that match a regular expression.
* Track any **statistics calculable for tabular data**, e.g., correlation changes between descriptors and model target.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-e999dc214d2a090d67b880e8ce3c6028c04492b3%2F2023109_course_module3.029-min.png?alt=media)

## Summing up

We introduced basic concepts and examples of text descriptors. You can calculate myriad descriptors, and we recommend collaborating with someone from your business team to determine text descriptors that fit your problem statement best.

While text descriptors are a powerful tool to analyze and monitor unstructured text data, we can use this approach only if the raw text is available. Other monitoring strategies may be required for scenarios where only embeddings are available.

Further reading: [Monitoring unstructured data for LLM and NLP with text descriptors](https://www.evidentlyai.com/blog/unstructured-data-monitoring)

Up next: monitoring embeddings drift.


# 3.4. Monitoring embeddings drift

Strategies for monitoring embedding drift using distance metrics, model-based drift detection, and share of drifted components.

{% embed url="<https://www.youtube.com/watch?v=0XtABbNYU7U&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=18>" %}

**Video 4**. [Monitoring embeddings drift](https://www.youtube.com/watch?v=0XtABbNYU7U\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=18), by Emeli Dral

## What are embeddings?

Embeddings are numerical representations of the input data. They transform raw data like text, images, videos, or music into numerical vectors in high-dimensional space. Embeddings are frequently used in machine learning for classification, regression, and ranking tasks with unstructured data.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-ff980def9eb2d937b6a05bf342a254178849e3da%2F2023109_course_module3.032-min.png?alt=media)

## Embedding drift detection methods

Since embeddings are numerical values, we can use many methods to monitor embedding distribution drift.

**Distance metrics** Each object in the dataset, represented as an embedding, is a numerical vector. Distance metrics allow calculating distances between these vectors. For example, you can use Euclidean distance or Cosine distance (to assess the angle between vectors). We can detect shifts in datasets by measuring the distance between centroids of reference data and current data.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a80babd58d0dbbb7614d636930519bb0d63169ac%2F2023109_course_module3.035-min.png?alt=media)

**Model-based drift detection** This approach uses embeddings to build a domain classifier that distinguishes between reference and current data. The idea is similar to model-based drift detection on raw text data: you get an estimation of the model's ability to distinguish between datasets. However, with embeddings, this approach has a limitation: you cannot use the information about the strongest features or best objects to determine the root cause/source of drift.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-23ad55a555b7ed2c642a2052fb6e501e256665bc%2F2023109_course_module3.036-min.png?alt=media)

**Share of drifted components** You can also use the share of drifted components to monitor embedding drift. This approach treats each embedding component independently and uses drift detection methods that can be applied to numerical values. For each component, drift size or score is assessed. You can then aggregate these individual scores into the number of drifted components or share of drifted components.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-da0704feb504937e8d08d7ad41fad40f2f74a161%2F2023109_course_module3.037-min.png?alt=media)

{% hint style="info" %}
**Further reading:** [Shift happens: we compared 5 methods to detect drift in ML embeddings](https://www.evidentlyai.com/blog/embedding-drift-detection).
{% endhint %}

The suggested blog might be especially interesting to those who work a lot of embeddings. It compares various methods to detect drift in embeddings – Euclidean distance, Cosine distance, Domain classifier, Share of drifted components, and Maximum mean discrepancy (MMD) – and evaluates each method against such criteria as computational speed, thresholds, PCA, and embedding model behavior.

Domain classifier can be a good default: it is comparably fast, PCA-agnostic, agnostic to the embedding model, and easy to interpret. However, we suggest you check out the blog to choose the right drift detection method for your specific use case. Here is a blog summary sneak peek:

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-17b2c66abdb6c338fa207502409cd24a44973034%2F2023109_course_module3.039-min.png?alt=media)

## Summing up

We discussed different strategies for monitoring embedding drift, including distance metrics, model-based drift detection, and share of drifted components. While using a domain classifier to detect embedding drift is a good default strategy, we suggest you evaluate other methods to choose the right one for your use case.

Further reading: [Shift happens: we compared 5 methods to detect drift in ML embeddings](https://www.evidentlyai.com/blog/embedding-drift-detection)

Up-next: code practice! We will apply ML monitoring strategies for unstructured data on real data to derive actionable metrics.


# 3.5. Monitoring text data \[CODE PRACTICE]

A code example walkthrough of unstructured data evaluations using the open-source Evidently Python library.

{% embed url="<https://www.youtube.com/watch?v=RIultWCjYXo&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=19>" %}

**Video 5**. [Monitoring text data \[CODE PRACTICE\]](https://www.youtube.com/watch?v=RIultWCjYXo\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=19), by Emeli Dral

In this video, we walk you through the code example of evaluations for unstructured data using the open-source [Evidently](https://github.com/evidentlyai/evidently) Python library.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module3/unstructured_data_code_practice.ipynb) to follow along.

**Links to docs:**

* [Text overview](https://docs.evidentlyai.com/presets/text-overview)
* [Embeddings](https://docs.evidentlyai.com/user-guide/customization/embeddings-drift-parameters)

**Outline:**\
[00:00](https://youtu.be/RIultWCjYXo?si=5s0_-fMduGKorqci) Import libraries and datasets\
[02:26](https://youtu.be/RIultWCjYXo?si=Vyrnq26avImqSUB6\&t=146) Prepare a multimodal dataset with raw text\
[08:15](https://youtu.be/RIultWCjYXo?si=hKrfvOBPZ3kFeisC\&t=495) Text data overview report\
[11:18](https://youtu.be/RIultWCjYXo?si=nws_RxLC2YsoiD1C\&t=678) Model-based text data drift detection\
[13:49](https://youtu.be/RIultWCjYXo?si=WlLgpbHHt2Bi-UIH\&t=829) Adding custom text descriptors to report\
[17:50](https://youtu.be/RIultWCjYXo?si=9KWXxjYqW4eaE97n\&t=1070) Custom report with descriptor drift detection\
[23:55](https://youtu.be/RIultWCjYXo?si=9lNBhLuipZrDK8zi\&t=1435) Embedding drift detection

That’s it! We covered three use cases for evaluating unstructured data – multimodal data with raw text, detecting drift in text data with descriptors, and embedding drift detection – and learned to derive actionable metrics.


# 3.6. Monitoring multimodal datasets

Strategies for monitoring data quality and data drift in multimodal datasets.

{% embed url="<https://www.youtube.com/watch?v=b0a1iMlHgEs&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF&index=20>" %}

**Video 6**. [Monitoring multimodal datasets](https://www.youtube.com/watch?v=b0a1iMlHgEs\&list=PL9omX6impEuOpTezeRF-M04BW3VfnPBRF\&index=20), by Emeli Dral

## What is a multimodal dataset?

Often, we don't only work with structured or unstructured data but a combination of both. Some common examples include product reviews, chats, support tickets, and emails. These applications may include unstructured data, e.g., text, and structured metadata like the region, device, product type, user category, etc.

Both structured and unstructured data provide valuable signals. Considering signals from both data types is essential to build comprehensive ML models.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-0ec6f165910b0cfe036656212081a34fb8ce24ba%2F2023109_course_module3.044-min.png?alt=media)

## Monitoring strategies for multi-modal data

We will cover three widely used strategies for monitoring multi-modal datasets.

**Strategy 1. Split and monitor independently.** The approach is straightforward – split the dataset by data type and monitor structured and unstructured data independently:

* Monitor structured data using descriptive statistics, share of missing values, distribution drift, correlation changes, etc.
* Use raw text data analysis or embedding monitoring for unstructured data.
* Combine monitoring results into a unified dashboard.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-05eb571a19ff16482bdfd4bbd39cc708f832c8ca%2F2023109_course_module3.046-min.png?alt=media)

**Strategy 2. A joint structured dataset.** This approach is based on turning unstructured data into structured by using descriptors:

* Generate descriptors for unstructured data (e.g., text properties) to represent it in a structured form.
* Combine these structured descriptors with existing metadata.
* Perform a comprehensive analysis of the combined structured data. You can check for missing values, distribution drift, correlation changes, outliers, etc.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-425cb46626d3d34b1c6bf53c90caecac18f81f12%2F2023109_course_module3.047-min.png?alt=media)

**Strategy 3. Generate embeddings.** As embeddings represent data as vectors in high-dimensional space, you can combine structured features with embeddings to create an expanded feature space. For instance, if you have 64 embeddings and three structured features, the combined space would be 67-dimensional. You can then apply various methods like share of drifted components, domain classifier, or distance-based metrics to this combined data.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-f9cb0cae755d9c1012552a5c30528eb95de79d43%2F2023109_course_module3.048-min.png?alt=media)

## Summing up

We discussed three strategies for monitoring data quality and data drift in multi-modal datasets. This concludes our module on ML monitoring for unstructured data. Here are some considerations to keep in mind:

* **If you have access to raw text data, do not ignore it.** Interpretability wins! Evaluating metrics on raw text can provide a deep understanding of changes and potential issues with text data.
* **If working with embeddings,** numerous methods are also available to detect drift.
* **When dealing with multimodal datasets,** you can split data by type, leverage text descriptors, or generate a joint embedding dataset, depending on the specific use case and available data.

## Enjoyed the content?

Star Evidently on GitHub to contribute back! This helps us create free, open-source tools and content for the community.

⭐️ [Star](https://github.com/evidentlyai/evidently) on GitHub!


# Module 4: Designing effective ML monitoring

This module reviews how to set up an ML monitoring system, considering the model risks, criticality, and deployment scenario.

In previous modules, we reviewed possible metrics and approaches to tracking the performance of ML models in production.

In this module, we will put it all together and review specific questions you might have when setting up an ML monitoring system for a particular model. We’ll do a deeper dive and cover:

* How to select and prioritize ML monitoring metrics.
* How and when to retrain ML models.
* How to choose a reference dataset.
* How to implement custom metrics in ML monitoring.
* How to choose an appropriate ML monitoring architecture.

At the end of this module, you will understand how to design an optimal approach to ML monitoring considering the model risks, criticality, and deployment scenario.


# 4.1. Logging for ML monitoring

What a good ML monitoring system is, and how to set up the logging architecture to capture metrics for further analysis.

{% embed url="<https://youtu.be/CtUsDcA3tB0?si=RNDR2uRZ7wc8NwxB>" %}

**Video 1**. [Logging for ML monitoring](https://youtu.be/CtUsDcA3tB0?si=RNDR2uRZ7wc8NwxB), by Emeli Dral

## What is a good ML monitoring system?

A good ML monitoring system consists of three key components:

* **Instrumentation** to ensure collection and computation of useful metrics for analyzing model behavior and resolving issues.
* **Alerting** to define unexpected model behavior through metrics and thresholds and design action policy.
* **Debugging** to provide engineers with context to understand model issues for faster resolution.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-1db6977872f57af17cf93615ec7572c1400a13a2%2F2023110_course_module4_fin.004-min.png?alt=media)

When it comes to ML monitoring setup, there is no “one size fits all.” Here are some factors that affect the ML monitoring architecture and choice of metrics:

* ML service implementation (online service vs. batch model).
* Environment stability.
* Feedback loop (immediate or delayed feedback).
* Team resources (capacity to implement and operate the ML monitoring system).
* Use case criticality.
* Scale and complexity of the ML system.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-d2870be1023431dc7df3f1c03594fd0ebb0e3b9e%2F2023110_course_module4_fin.005-min.png?alt=media)

## Logging and instrumentation

ML monitoring starts with logging. Before talking about metrics, you need to implement a way to collect the data for analysis.

**Step 1. Capture service (event) logs**

Capturing service logs is a must-have for any production service, as it helps to monitor and debug service health. You may record different types of events that happen in your service. One will be the prediction event when the service gets the input data and returns the output.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-f4e138a85132f6b33532527b62f6013084cb4c37%2F2023110_course_module4_fin.008-min.png?alt=media)

**Step 2. Capture prediction logs**

When you record the prediction event, make sure to log all prediction-related information, including model input data, model output, and ground truth, if available.

These prediction logs are the key input for ML model quality monitoring. You also need them for model retraining, debugging, and audits.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-b65463b1347a733f65f0fc36b5dedc8cfd60bd82%2F2023110_course_module4_fin.009-min.png?alt=media)

**Step 3. Log ML monitoring metrics**

Logging architecture heavily depends on how you deploy your models.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-8cd0d02830521d3f81625d97ea0edd57f95b9c9b%2F2023110_course_module4_fin.012-min.png?alt=media)

Typically, it involves a **prediction store** where prediction data is recorded. It requires long-term secure storage with features like backups.

Once you set up this prediction logging, here are two main approaches to **ML monitoring implementation**:

* **ML monitoring service** can pull data from the prediction store – or data can be pushed directly from the ML service – to compute monitoring metrics.
* **Monitoring jobs** can be operated with the help of a pipeline manager and load data from the prediction store to compute monitoring metrics.

The next element is a **monitoring dashboard**:

* A **metric store** is created to store computed metrics. It needs quick querying capabilities for efficient dashboard interactions.
* The monitoring dashboard uses this metric store as the **data source** to visualize calculated metrics.

**For smaller datasets**, connecting the monitoring dashboard directly to the prediction store can suffice. It also works well if you run **ad-hoc or scheduled reports**.

## Summing up

We discussed setting up the logging architecture to capture useful metrics for further analysis. Next, we will cover what exactly to log.


# 4.2. How to prioritize ML monitoring metrics

ML monitoring depth, metrics to collect when building a monitoring system, and how to prioritize them.

{% embed url="<https://youtu.be/jCXO4uuMHbs?si=9ss6_nK2bbLGq8Ph>" %}

**Video 2**. [How to prioritize ML monitoring metrics](https://youtu.be/jCXO4uuMHbs?si=9ss6_nK2bbLGq8Ph), by Emeli Dral

## Which metrics to collect?

What exactly to log and which metrics to collect depend on your primary monitoring goal.

For example, if you focus on **alerting**:

* Log a few metrics to be able to detect when things go wrong.
* Your main goal is to detect and alert on any issues quickly.

If you focus on **debugging and analysis**:

* Log additional context to understand and cover potential problems.
* Include not only main metrics but also some descriptive statistics to cover unknown unknowns.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-ded9c4dd9b7b7c108bbef81e2814d27ce854c9ab%2F2023110_course_module4_fin.018-min.png?alt=media)

## Metric hierarchy

Let us introduce a hierarchy of metrics you may collect. The hierarchy consists of five layers:

* **Symptoms** are key metrics that help identify model issues, e.g., 1-day accuracy or business KPIs. They are used to build alerts.
* **Summary metrics** provide additional context. It can be metrics like prediction volume, true positives, true negatives, share of missing data, etc.
* **Input/output and performance profiling** give you more details, such as per-column descriptive statistics, distributions, segmented model performance, etc.
* **Debugging and analytics data** include metrics for deeper analysis, like correlations and explainability.
* **Raw data prediction logs** can be queried to calculate any custom metrics or build custom visualizations.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-bc732c08e18c6c6490a56a748f7e91b1b3b27221%2F2023110_course_module4_fin.019-min.png?alt=media)

## Monitoring depth

Depending on your specific circumstances, you can choose minimalistic symptom-based monitoring: log less data and only look at a few metrics.

**Minimalistic monitoring** is totally fine if you:

* Have just **a few ML models** in production,
* Work with **smaller datasets**,
* Deal with **less critical** use cases,
* Build your monitoring for **technical users** only,
* Can **easily access predictions** for any additional analysis,
* If the **prediction frequency** is low.

However, you might need more **in-depth monitoring** if you:

* Have **multiple models in production** and want to standardize debugging and model support,
* Work with **larger datasets**, and hence, it is expensive to query logs,
* Deal with **critical use cases** where the cost of error is high,
* Build your monitoring system for **business stakeholders**.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-3dec601925e5af94802c3b6720e83897cd05d95f%2F2023110_course_module4_fin.020-min.png?alt=media)

**Is there a good default?** Monitoring is use-case specific. However, for each new model, you can start with a “middle ground” metric composition:

* Pick a few **key quality indicators** to use as issue symptoms to alert on.
* Add **per-column data summaries** and **performance summaries** to simplify immediate troubleshooting.

This can be a good go-to solution until you develop a monitoring approach that fits your specific use case.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-7e664475aa8f41758b2619f4bbc2c339a1a9a8c3%2F2023110_course_module4_fin.021-min.png?alt=media)

## Metric prioritization

Here is how you can prioritize metrics for your initial setup:

* Service health
* Model performance
* Data quality and integrity
* Data and concept drift

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-b02521987311461b669bd54b72df7bbdb4130aba%2F2023110_course_module4_fin.023-min.png?alt=media)

**1. Service health**

Service health is a must-to-track as it forms a basis for all other metrics. It answers the question "Does the service work?" To track the service's health, you can use standard software performance metrics like latency, error rate, memory usage, disk usage, etc.

An important thing to add is monitoring the number of model predictions. This is especially relevant if you have fallback solutions such as alternative models or rule-based systems you use when the model does not respond or the incoming data has low quality. You need to know how often each model is used.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-0667179a4d357638fd8dc98a66e1665326a71a8b%2F2023110_course_module4_fin.025-min.png?alt=media)

**2. Model performance**

This set of metrics answers the question "How does the service perform?" and “Did anything break?”

**Monitoring.** Depending on your use case, you can pick the best available signal to alert on: **direct** or **proxy metrics**.

* **Direct**. You can measure business KPIs or direct model quality (accuracy, RMSE, etc.).
* **Proxy**. If the ground truth is not available, you can switch to proxy metrics, e.g., use heuristics, check that output range compliance, or monitor output distribution drift.

Make sure that your chosen KPIs are something you want to act upon if a certain threshold is reached.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-8adc52a756ea672439f3c9b3e3a859217df771fc%2F2023110_course_module4_fin.028-min.png?alt=media)

**Debugging**. To simplify troubleshooting, you can also capture additional statistics like distribution output shape. For example, if the prediction drift is detected, you can see how the distribution of the model predictions has changed. You can also log data on model errors to analyze error buckets and error distribution.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-5adc628fe02aed1828237533bd9cc81c2ed36dd3%2F2023110_course_module4_fin.029-min.png?alt=media)

**3. Data quality and data integrity**

If model performance issues are detected, data quality and integrity metrics help to understand where it has broken and whether you can trust your data.

**Monitoring**. There are two methods you can use to analyze data quality and data integrity:

* **Track data health metrics**, such as the share of missing values, the share of duplicates, and the share of features out of range.
* **Run test suites** on top of data health metrics. You can periodically run scans with a large number of checks and track the number of failed tests.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-45fb1e5b250b69c203bfa404087c986b61cf6f46%2F2023110_course_module4_fin.033-min.png?alt=media)

**Debugging**. To log additional data for debugging purposes, you can capture dataset or column statistics – count, mean, standard deviation, min, max, etc.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-682da5d3d825075fc7d937aaf841be1b8b233494%2F2023110_course_module4_fin.034-min.png?alt=media)

**4. Data and concept drift**

This block of metrics helps to determine whether the model is still relevant.

**Monitoring**. To simplify the alerting and avoid false alerts, you can look at the following:

* Distribution drift in the most important features only.
* Prediction drift.
* Overalldataset drift instead of drift in the individual features.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-fc7dc5967bf7be0fb7420277eea305d3a8515c98%2F2023110_course_module4_fin.036-min.png?alt=media)

**Debugging**. You can log feature distributions to compare them over time.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-d9782404cd3b9b176f4e023e7ca4c34de5d144e2%2F2023110_course_module4_fin.037-min.png?alt=media)

## Comprehensive monitoring

Depending on the problem statement and model usage scenario, you can introduce more comprehensive monitoring metrics:

* **Performance by segment**. It can be especially useful if you deal with a diverse audience or complex object structures and want to monitor them separately.
* **Model bias and fairness**. These metrics are crucial for sensitive domain areas like healthcare.
* **Outliers**. Monitoring for outliers is vital when individual errors are costly.
* **Explainability**. Explainability is important when users need to understand model decisions/outputs.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-74f9c035540ac5041b47c759adf7807aa7296634%2F2023110_course_module4_fin.038-min.png?alt=media)

## Summing up

We discussed which metrics to collect when building a monitoring system and how to prioritize them. Here are some must-haves to keep in mind:

* Always monitor service health metrics.
* Always store prediction logs, including inputs and outputs.
* Limit alerting to key metrics or symptoms.
* Define monitoring depth based on your specific needs.
* When in doubt, go for the middle ground: select several metric groups – e.g., model quality, data quality, and data drift – and define the best signals for each group.

Up next: dive into the retraining system and its connection to model performance and monitoring.


# 4.3. When to retrain machine learning models

Scheduled and trigger-based retraining and what to consider when making the retaining decision.

{% embed url="<https://youtu.be/oqyyVp-t5A8?si=u-DeeyxCx9fBjs1>\_" %}

**Video 3**. [When to retrain machine learning models](https://youtu.be/oqyyVp-t5A8?si=u-DeeyxCx9fBjs1_), by Emeli Dral

## Why retrain ML models?

ML model quality degrades with time. It happens due to changes in the environment models operate within. The goal of ML monitoring is to help detect when the ML model’s quality starts to degrade, intervene, and get the ML service back on track.

The good news is that new data is accumulated during model use. You can use this data to retrain and update your model.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-480dde4719f176c62d2ffed5e8fc6506b850888e%2F2023110_course_module4_fin.042-min.png?alt=media)

## Model retraining strategies

We will cover the two most widely used retraining strategies: scheduled retraining and trigger-based retraining.

**Scheduled retraining** This approach is straightforward: define a retraining schedule – daily, weekly, or monthly – and stick to it. The optimal schedule depends on the use case and data stability. This approach is a good enough solution to start with.

**Pros**:

* Simplicity in execution; no need to overthink. **Cons**:
* Resource-intensive.
* May require complex infrastructure for model evaluation and deployment automation.
* More expensive support for automated retraining and deployment.
* Sometimes, it can make things worse if you don’t have a solid procedure for model evaluation in place. You risk deploying a model that is worse than the current one.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-153e23375f616e097c7ba070dbcbc3fef7e1f6e5%2F2023110_course_module4_fin.045-min.png?alt=media)

**Pro-tip for scheduled retraining**: use historical data to determine the rate of model decay and the volume of new data required for effective retraining. For example, you can get a training set from your historical data and train a model on top of this dataset. Then, you can start experimenting: apply this model to new batches of data with a certain time step – daily, weekly, monthly – to measure how the model performs on the new data and define when its quality starts to degrade. You can also run checks on historical data to define:

* Will more data improve model quality?
* How quickly does the quality degrade?
* How much data is needed to retrain the model?
* Should you drop the old data when retraining?

**Important note**: you need labels to do it. If feedback/ground truth is not available yet, it makes sense to send data for labeling before you start experimenting with historical data.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-4041383cff1d7a4ad60a8b04086da528cc08517f%2F2023110_course_module4_fin.046-min.png?alt=media)

{% hint style="info" %}
**Further reading:** [To retrain, or not to retrain? Let's get analytical about ML model updates](https://www.evidentlyai.com/blog/retrain-or-not-retrain).
{% endhint %}

**Trigger-based retraining** To use this strategy, you need to pick a key metric – e.g., daily accuracy – and retrain the model if this metric breaks a certain threshold. For example, the retraining trigger may be “retrain the model if daily accuracy drops under 0.9.”

**Pros**:

* Allows retraining the model automatically.
* Works well for costly updates and complex approvals, as retraining is based on specific performance metric drops. **Cons**:
* Requires a robust monitoring system.
* Risk of missing early quality drops.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-486d002a1cf0fcdada2372ff77211279c0f51aab%2F2023110_course_module4_fin.048-min.png?alt=media)

## Model retraining tradeoffs

The model retraining process isn't just about updating the model. It includes generating and checking the quality of the new dataset, evaluating the model, and updating the service – all of these come with associated costs.

When it comes to retraining, there is always a **tradeoff** between the benefits of improved/sustained model performance and the resources you invest in retraining, such as time, computational and human resources, process automation, and support.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-483041d401fe1d689a0681fb9bf8f5a21a1a8aae%2F2023110_course_module4_fin.050-min.png?alt=media)

## Thinking through the retraining decision

Here are some considerations when deciding whether to retrain the model or not:

**Be pragmatic**. Develop a strategy considering available actions, service properties, resources, model criticality, and the cost of errors. Here is an example of a decision-making logic you can follow:

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-2477f565c9ef4d64e6528ae16a8a7bd2afdcdec3%2F2023110_course_module4_fin.054-min.png?alt=media)

Let’s look at each of the **steps of this decision-making logic** in more detail.

**Check for model performance**. To define whether the model quality has dropped, compare the model performance against a baseline. You should check if the drop is real: it is important to distinguish between the case when the quality dropped for a couple of objects and the case when the quality dropped on top of the whole batch of data.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-793e61b28146a99934eaf84fd558ed677d960f4a%2F2023110_course_module4_fin.055-min.png?alt=media)

**Investigate data quality issues**. If something is wrong with the model, chances are there is something wrong with the data. To investigate potential data quality issues, you can run tests to check for missing values, duplicates, out-of-range values, etc. Remember to **fix the data quality first** and then confirm if retraining is needed.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-44ee0b21800cdef1d4a2e51601c5df850d34ab82%2F2023110_course_module4_fin.056-min.png?alt=media)

**Investigate data drift**. If the data quality is OK, it can be the data shifts that cause model performance issues. To detect them, check for distribution drift in model target behavior and input data.

If there are no shifts in the data, look for **technical bugs**. If you detect data drift, consider both **retraining** and **switching to alternatives** – rule-based systems, other models, or manual processes.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-4f479e7de394e0d94e2330a31244047430ab5ea7%2F2023110_course_module4_fin.057-min.png?alt=media)

**Decide on retraining if sufficient data is available**. Sometimes, you may have sufficient data to detect drift but not enough to retrain the model. It is useful to come up with a heuristic to define the minimum volume of data you need to retrain the model.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-bd94ba130c93c061594b278baedc082701fb7363%2F2023110_course_module4_fin.058-min.png?alt=media)

**Check your alternatives**. You can consider switching to an alternative decision-making process, like the previous model version, manual processing, non-ML models, or use heuristics.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-6db7a8886563845b5c896bc334481e874b1feb73%2F2023110_course_module4_fin.059-min.png?alt=media)

**Evaluate model quality before rollout**. You can compare how your alternatives perform on a "golden dataset" which:

* Contains all relevant segments/classes,
* Includes known corner cases or test scenarios,
* Represents the long-term trends, not only recent data.

**The golden dataset should be curated** to reflect assumptions about the data and incorporate domain knowledge.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-450a0ab47c8b23fce1f9e1c586ec745ea8d099c9%2F2023110_course_module4_fin.060-min.png?alt=media)

## Summing up

We discussed two main strategies for model retraining: scheduled retraining and trigger-based retraining. Scheduled retraining is a good enough approach but rarely optimal, as you might miss the true decay or waste resources.

How to do better:

* Make sure you evaluate the models before roll-out.
* Monitor the production model quality and data.
* Test your retraining strategy on historical data to make informed choices.
* Tune and automate retraining the approach, but keep the human in the loop.

Further reading: [To retrain, or not to retrain? Let's get analytical about ML model updates](https://www.evidentlyai.com/blog/retrain-or-not-retrain)

Up next: we will discuss how to create and curate a reference dataset.


# 4.4. How to choose a reference dataset in ML monitoring

What a reference dataset is in ML monitoring, how to choose one for drift detection, and when to use multiple references.

{% embed url="<https://youtu.be/42J-C4WmkZc?si=Av1gwZXAkBZXDT70>" %}

**Video 4**. [How to choose a reference dataset in ML monitoring](https://youtu.be/42J-C4WmkZc?si=Av1gwZXAkBZXDT70), by Emeli Dral

## Why use a reference dataset?

There are two main uses for a reference dataset.

1. You can use it **to derive test conditions automatically**, saving time and effort in setting up tests manually.

For example, you can use a reference dataset to generate conditions for data quality checks (to track feature ranges, share of missing values, etc.) and model quality checks (to keep tabs on metrics like precision and accuracy) by passing a previous batch of data as a reference.

2. You can use a reference dataset **as a baseline to detect data and prediction drift** in production by comparing new data distributions against reference distributions.

Reference dataset also can be used to **detect training-serving skew** as it provides a baseline to detect changes between training and production data.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-7029dd16b0ad0af9deca04c5a320ec27665d0b76%2F2023110_course_module4_fin.064-min.png?alt=media)

## What makes a good reference dataset?

Characteristics of a **good reference dataset**:

* Reflects realistic data patterns, including cycles and seasonality.
* Contains a large enough sample to derive meaningful statistics.
* Includes realistic scenarios (e.g., sensor outages) to validate against new data.

What a **reference dataset is not**:

* It is **not the same as a training dataset**. You can sometimes choose training data to be your reference, but they are not synonymous.
* It is **not a “golden dataset,”** which serves a different purpose.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-a85c878c834aa50fedf7d3506b859e7ba69461ea%2F2023110_course_module4_fin.065-min.png?alt=media)

You always need a reference dataset if your goal is to compare distributions to detect data or prediction drift (e.g., using metrics like Wasserstein distance).

However, having a reference dataset is not a must:

* You can run one-sample statistical tests that don't require a comparison of distributions.
* For most types of checks, you can manually specify test conditions, such as min-max feature ranges. This works well if you have a limited set of data with known expected behaviors.

However, using a reference dataset is a great hack to automate generating test conditions!

## Using training data as a reference

Using training data as a reference can be acceptable in specific contexts but is generally not recommended due to pre-processing and potential biases.

**If the training data is all you have**, it is OK to use it for the following types of checks:

* To derive feature types and data schema.
* To derive feature ranges (num) and value lists (cat).
* To derive feature correlations.
* To detect training-serving skew.

**It is less optimal for**:

* Generating expectations about model quality.
* Deriving data on e.g., share of nulls.
* Using it as a baseline for drift detection.

You can consider using hold-out validation data or previous batches of data instead.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-f04c21f9b0c8a396c99a6cb8d8f3af5015bbd31c%2F2023110_course_module4_fin.068-min.png?alt=media)

## Reference dataset for drift detection

When choosing a reference dataset for drift detection, make sure to pick a representative dataset that captures typical distributions and variations in the data. You can use historical data to decide on the appropriate windows; for example, you can compare the data using monthly, weekly, or daily windows.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-7d6d63fb41a368cf6fd4f9ed4f6bd6a98db808a5%2F2023110_course_module4_fin.069-min.png?alt=media)

You should make the following decisions:

* **What do you compare against?** You can use training data (generally not recommended), validation, and previous production batches.
* **What batch size to use?** You need to determine the size of current and reference datasets for effective comparison – 1 day, 1 week, 1 year, 1000 objects, etc.
* **How to update reference data?** You can have a static reference (e.g., which you update once a month) or shift the reference data dynamically (e.g., sliding window approach).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-6cf69e34d5b085a3b5bb9703e341eb89150b228d%2F2023110_course_module4_fin.071-min.png?alt=media)

Analyzing historical data can help determine the most effective reference data strategy.

{% hint style="info" %}
**Further reading:** [How to detect, evaluate and visualize historical drifts in the data](https://www.evidentlyai.com/blog/tutorial-3-historical-data-drift).
{% endhint %}

**Multiple references**. It often makes sense to use multiple reference datasets. For example, you can have multiple comparison windows to capture seasonality and cyclic trends:

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-513b5f02a64f542b3fddf22a46f814fd69173c52%2F2023110_course_module4_fin.073-min.png?alt=media)

**Sampling vs. entire dataset**. If you have large datasets, you can consider using sampling, for example, random or stratified sampling.

* **Sampling is great for drift detection**. In fact, all statistical tests were made to work with samples! It can save you computational resources and allow for faster results calculation. If you look to detect a statistical distribution shift in the overall dataset, sampling is totally fine.
* **For detecting data quality anomalies, full datasets are preferable**. If you look for data quality issues – e.g., individual outliers or duplicates – sampling can disguise them.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-8a9c309c87ea9f86264f2f21f1801b3fa43e1213%2F2023110_course_module4_fin.074-min.png?alt=media)

## Summing up

* There is no “universal” reference dataset. It should be tailored to the specific use case and expectations of similarity to current data.
* Hold-out validation data is preferred over training data for creating reference datasets, especially for drift detection. Use training data only if there is nothing else.
* It is crucial to account for seasonality and historical patterns when choosing the reference dataset to ensure that it accurately represents the expected variations in data.
* Historical data is a valuable resource for informing reference dataset selection.

Further reading: [How to detect, evaluate and visualize historical drifts in the data](https://www.evidentlyai.com/blog/tutorial-3-historical-data-drift)

Up next: custom metrics for ML monitoring.


# 4.5. Custom metrics in ML monitoring

Types of custom metrics. Business or product metrics, domain-specific metrics, and weighted metrics.

{% embed url="<https://youtu.be/PrFuzKLM66I?si=68EF7tepIyXxyMig>" %}

**Video 5**. [Custom metrics in ML monitoring](https://youtu.be/PrFuzKLM66I?si=68EF7tepIyXxyMig), by Emeli Dral

## Types of custom metrics

While there is no strict division between “standard” and “custom” metrics, there is some consensus on evaluating, for example, classification model quality using metrics like precision and recall. They are fairly “standard.”

However, you often need to implement “custom” metrics to reflect specific aspects of model performance. They typically refer to business objectives or domain requirements and help capture the impact of an ML model within its operational context.

Here are some examples.

**Business and product KPIs (or proxies)**. These metrics are aligned with key performance indicators that reflect the business goals and product performance.

**Examples include**:

* Manufacturing optimization: raw materials saved.
* Chatbots: number of successful chat completions.
* Fraud detection: number of detected fraud cases over $50,000.
* Recommender systems: share of recommendation blocks without clicks.

We recommend **consulting with business stakeholders** even before building the model. They may suggest valuable KPIs, heuristics, and metrics that could be monitored even during the experimentation phase.

When direct measurement of a KPI is not possible, consider **approximating the model impact**. For example, you can assign an average “cost” to specific types of model errors based on domain knowledge.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-c43483ff85587587f2fefebc903b8a6ce124840b%2F2023110_course_module4_fin.078-min.png?alt=media)

**Domain-specific ML metrics**. These are metrics that are commonly used in specific domains and industries.

**Examples include**:

* Churn prediction in telecommunications: lift metrics.
* Recommender systems: serendipity or novelty metrics.
* Healthcare: fairness metrics.
* Speech recognition: word error rate.
* Medical imaging: Jaccard index.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-2cd99c013403637cdcdc89f9755427192bcd05d3%2F2023110_course_module4_fin.079-min.png?alt=media)

**Weighted or aggregated metrics**. Sometimes, you can design custom metrics as a “weighted” variation of other metrics. For example, you can adjust them to account for the importance of certain features or classes in your data.

**Examples include**:

* Data drift weighted by feature importance.
* Measuring specific recommender system biases, for example, based on product popularity, price, or product group.
* In unbalanced classification problems, you can weigh precision and recall by class or by specific important user groups, such as based on the estimated user Lifetime Value (LTV).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-9a29f0a36859502b8aed33d4c9a906e2c97969e5%2F2023110_course_module4_fin.080-min.png?alt=media)

## Summing up

There is no need to invent “custom” metrics just for the sake of it. However, you might want to implement them to:

* better reflect important model qualities,
* estimate the business impact of the model,
* add metrics useful for product and business stakeholders and accepted within the domain.

Up next: optional code practice to create and implement a custom quality metric in the Evidently Python library.


# 4.6. Implementing custom metrics in Evidently \[OPTIONAL]

A code example walkthrough of creating a custom metric using the Evidently Python library.

{% embed url="<https://youtu.be/uEyoP-sPhyc?si=7hwr4LaJIeBZ-YLD>" %}

**Video 6**. [Implementing custom metrics in Evidently \[OPTIONAL, CODE PRACTICE\]](https://youtu.be/uEyoP-sPhyc?si=7hwr4LaJIeBZ-YLD), by Emeli Dral

This is an optional code practice video. It is useful when you already have experience using the Evidently Python library and are familiar with the existing Metrics and Tests. If you are new - check out the next module for an end-to-end example!

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module4/custom_metric_practice.ipynb) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=uEyoP-sPhyc\&t=0s) Introduction\
[00:37](https://www.youtube.com/watch?v=uEyoP-sPhyc\&t=37s) Imports\
[01:54](https://www.youtube.com/watch?v=uEyoP-sPhyc\&t=114s) Understanding the structure of Metrics and Tests\
[05:11](https://www.youtube.com/watch?v=uEyoP-sPhyc\&t=311s) Create a dummy custom metric\
[12:17](https://www.youtube.com/watch?v=uEyoP-sPhyc\&t=737s) Apply a dummy metric on toy data\
[14:00](https://www.youtube.com/watch?v=uEyoP-sPhyc\&t=840s) Create a more complicated metric: Mean by Category\
[26:25](https://www.youtube.com/watch?v=uEyoP-sPhyc\&t=1585s) Apply a new metric on toy data


# 4.7. How to choose the ML monitoring deployment architecture

ML monitoring architectures and choosing the right architecture for your use case.

{% embed url="<https://youtu.be/Q1NUCDZFRbU?si=26GhKBdhFAIzxBgi>" %}

**Video 7**. [How to choose the ML monitoring deployment architecture](https://youtu.be/Q1NUCDZFRbU?si=26GhKBdhFAIzxBgi), by Emeli Dral

There are alternative backends for machine learning monitoring architecture.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-92e934d1fc13ad734ee7b58f99e0316f4f33701d%2F2023110_course_module4_fin.086-min.png?alt=media)

## Ad-hoc reporting

**Ad-hoc reporting** is a viable option when you've recently deployed a machine learning system, and do not have alternative monitoring systems.

* It has **low engineering overhead**: you can use familiar tools like Jupyter notebooks, Python scripts, or R scripts.
* It is **suitable for initial exploration** of data and model quality and shaping expectations about model performance, but is not a long-term monitoring solution.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-0a0e3e109df5f5135de668a2718c827fb72fbea9%2F2023110_course_module4_fin.087-min.png?alt=media)

## Batch monitoring

**Batch ML monitoring** is a reliable and stable approach. It is suitable for both machine learning pipelines and services.

To implement batch monitoring, you need a workflow orchestration tool like Airflow or Kubeflow, and tools for calculating metrics and tests, such as Evidently.

**Pros**:

* Works well for both ML models implemented as batch pipelines and ML services.
* It is fairly simple to run monitoring jobs, especially if you already have a workflow orchestrator in place.
* You can use the same tools you use to run model training jobs during the experimental and validation phases of a machine learning lifecycle.
* You can combine immediate monitoring (e.g., data quality checks) and metrics dependent on ground truth (trigger-based calculations).

**Cons**:

* It is not real-time. There are some delays in metric computation due to additional resources required for running the infrastructure.
* It might be complex if you don't have an existing orchestrator; setting up one can be resource-intensive.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-df81a56e5079c443cccf2d19c516e3d3111a4659%2F2023110_course_module4_fin.088-min.png?alt=media)

## Near real-time (streaming) monitoring

**Near real-time ML monitoring** architecture is suitable when you serve models as APIs and want to detect issues close to real-time. In this case, you push data from the machine learning service to the monitoring system.

You will need optimal storage solutions for time series data like Prometheus or Clickhouse, and tools like Grafana or Evidently for dashboarding and alerting.

**Pros**:

* Works for models deployed as an ML service as opposed to batch jobs.
* Suitable for scenarios when you need an immediate reaction to issues like missing data or outliers.

**Cons**:

* High operational costs. Make sure you have the resources to maintain an additional monitoring service.
* Potentially double effort. You will often still need to deal with delayed ground truth feedback and run batch monitoring jobs to calculate these metrics.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-e3eb55d3f4646f28416d8a35fed26c839f7d099d%2F2023110_course_module4_fin.089-min.png?alt=media)

**Custom monitoring backend**. You can also combine near real-time and batch monitoring.

For example, you can combine:

* **Real-time checks**. You can send the data available at serving time directly from the ML service to an ML monitoring system to run input and model output checks and to generate alerts.
* **Monitoring jobs**. For delayed ground truth or more complex checks, you can run monitoring jobs over prediction logs on a trigger or a schedule.
* **Dashboarding tool**. You can log all results to the same metric storage system and get a single dashboard with panels for batch and real-time checks.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-dec026f0d1a039b39278a8ae4a81b7f8f69c8e61%2F2023110_course_module4_fin.090-min.png?alt=media)

## A case for batch ML monitoring

Let’s go through the possible logic of choosing the ML monitoring architecture.

First, let’s contrast it to **traditional software health monitoring**. You can typically implement additional service endpoints for metrics. Then, you can use tools like Prometheus to pull the metrics from these endpoints and store high-frequency time series data. You can add alerting and dashboard tools that rely on these metrics as a data source.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-f193ea5a9ada4192b73184ba66ee743c3ead6f5f%2F2023110_course_module4_fin.092-min.png?alt=media)

However, integrating ML metrics into this same setup isn't as simple. Here is why:

* **Complex metrics**. Software metrics are usually more straightforward in terms of computation. You can run simple aggregations over data points like response times and memory usage. Some ML-related metrics (like the number of rows or missing values) are similar. But others, like model quality or statistical tests, involve more complex calculations.
* **Delayed feedback**. Model quality metrics like precision, recall or accuracy typically depend on delayed data. You cannot compute them at serving time and must wait for the labels. Once you calculate them, you must “backfill” time series data for the past period, since the moment you compute metrics is not the moment they refer to.
* **Reference dataset**. For checks like data and prediction drift, you must also pass a batch of data you are comparing against. This does not easily fit into traditional software architecture.

ML model monitoring may require additional components:

* **Metric calculation pipelines**. If you run metric computation as jobs, you can use the appropriate backend for complex evaluations, for example, not just a SQL-like query. You can run complex evaluations like statistical drift and behavioral tests.
* **Run several different pipelines**. You can split metrics into separate pipelines. Some will run on a schedule (for metrics you can compute immediately) and others will be triggered by events like receiving new labeled data.
* **Passing the reference data**. You can implement complex pipelines that would involve querying the reference data, loading it, and comparing it against the current data batch.

**Example**: you can cover the whole model lifecycle with batch checks and monitoring jobs.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-74f20b55f0039d93721b499ea3733e091410eb08%2F2023110_course_module4_fin.102-min.png?alt=media)

You can still combine this approach with traditional software monitoring system architecture. Once you implement a different metric computation backend for ML metrics, you can store the results in a metric storage and use it as a data source for your dashboarding system to visualize machine learning-related metrics.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-7143101b68d3035b8598312e89b90ad21f320b9e%2F2023110_course_module4_fin.103-min.png?alt=media)

You can add a few ML-related metrics to an existing dashboard or create a separate ML monitoring dashboard.

## Summing up

We discussed the differences between different ML monitoring architectures. Here are some takeaways:

* Choose the ML architecture that matches your available resources, risk mitigation needs, and the complexity of your machine learning model.
* Even if you deploy a model as a service, consider batch ML monitoring. It is a more lightweight option, especially if you have a workflow orchestrator in place. It can handle complex evaluation scenarios.

## Enjoyed the content?

Star Evidently on GitHub to contribute back! This helps us create free, open-source tools and content for the community.

⭐️ [Star](https://github.com/evidentlyai/evidently) on GitHub!


# Module 5: ML pipelines validation and testing

This code-focused module demonstrates how to deploy an end-to-end pipeline for data and ML model quality checks.

In previous modules, we covered what ML monitoring is, which metrics and tests to use, and what to consider in ML monitoring design. Now, let’s get to practice! This is a code-focused module.

We will apply the learnings and **implement data and model quality tests as part of a pipeline**. If you deal with batch models, such test-based monitoring can often cover all your needs. For online models, this can be a part of your setup. You can run batch checks when you get labeled data or retain the models.

We will go through an **end-to-end pipeline using a toy dataset**. We will train a model and design tests for data and model quality using Evidently. We will also explore how to automate the data pipeline testing using tools like Airflow, Prefect, and Mlflow.


# 5.1. Introduction to data and ML pipeline testing

A brief introduction to different types of tests and testing conditions and how to incorporate them in data and ML pipelines.

{% embed url="<https://youtu.be/v7YW97YJ5DA?si=mEzSmfoTjDFxHF28>" %}

**Video 1**. [Introduction to data and ML pipeline testing](https://youtu.be/v7YW97YJ5DA?si=mEzSmfoTjDFxHF28), by Emeli Dral

## When to perform testing

ML lifecycle involves various steps that require testing to ensure that our ML models function properly. Critical areas to test include the following steps of the ML lifecycle:

* During the feature engineering stage: testing the input data quality as it affects the whole pipeline.
* During model training (or retraining): model quality checks.
* During model serving: validating incoming data and model outputs.
* During performance monitoring: continuously testing the model quality to detect and resolve potential issues.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-eda2cea166120fe65cbb320435e8a8be7aee3ef3%2F202310_module5_fin.007-min.png?alt=media)

## How to perform testing

There are different types of checks you can use to test data and ML pipelines:

**Individual tests** A test is a **metric** with a condition. You can perform a certain evaluation or measurement on top of a data batch and compare it against a threshold or expectation. You can formulate almost anything as a test: assertions on feature values, expectations about model quality on a specific segment, etc. Whatever you can measure, you can design as a test.

Tests can be **column-level** (when metrics are calculated for a specific feature or column) or **dataset-level** (in this case, you calculate metrics for the whole dataset).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-8df345b0c9d6cce9c91c72746d9b989cef17ba0d%2F202310_module5_fin.010-min.png?alt=media)

**Test suites** Individual tests can be grouped into test suites. For each test in a test suite, you can define test criticality and set alerting conditions: for example, based on the number of failed critical tests.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-df6841c8eeede4c8db561c0528763b0c215447c5%2F202310_module5_fin.011-min.png?alt=media)

When you create a test, you must define the test conditions. There are two main strategies you can use for establishing test conditions:

**Reference-based conditions** You can use a reference dataset to derive conditions automatically rather than set conditions manually for each individual test. This is a great option for certain types of checks, such as testing column types (which are easy to derive from a reference example) and for ad hoc testing such as when you import a new batch of data and can immediately visually explore the test results. However, be careful when designing alerting, as auto-generated test conditions are not perfect and may be prone to false alerts or missed issues.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-c1d9041bf27885d2d2ae7b8821abc02310694c3f%2F202310_module5_fin.013-min.png?alt=media)

**Manually defined conditions** With this approach, you specify conditions for each test manually. This method does not require additional data and can be great for encoding specific conditions based on domain expertise.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-29fe83009306f9025581917ed3652e1265779b11%2F202310_module5_fin.015-min.png?alt=media)

You can also combine reference-based and manual conditions. For example, you can manually pass conditions for specific features and use a reference dataset to define test conditions for the rest of your dataset. Combining these approaches is possible with tools like [Evidently](https://github.com/evidentlyai/evidently).

## Test automation

If you want to test your data and ML models continuously, switching from ad-hoc checks to automated testing is a good idea. You can use workflow managers like Airflow, Kubeflow, or Prefect to automate testing as part of the ML pipeline. If you run your ML model in batch, just add a testing step to your pipeline.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-e404f029a9d358cc643d6195d940267044fe096d%2F202310_module5_fin.016-min.png?alt=media)

## Recording test results

If you already use logging tools like MLflow, you can use them to log test results as well. Evidently also offers a monitoring dashboard where you can visualize individual and aggregate test results to track them in time.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-111c3f7ee4eb3331603fd1120f28f3e3ed270c35%2F202310_module5_fin.017-min.png?alt=media)

## Example use case

The practical part of the module involves applying data and model quality tests on a toy dataset. Using the Bank marketing dataset, we will predict subscription outcomes from a marketing campaign. Data source: [bank marketing](https://archive.ics.uci.edu/dataset/222/bank+marketing).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-4b3a110443932cad41ef2ff52757062ff6cd33c8%2F202310_module5_fin.019-min.png?alt=media)

You will design training and prediction pipelines as part of the code practice. For the **training pipeline**, you will prepare data, calculate features, do model training and scoring, and incorporate data, feature and model quality checks.

For the **prediction pipeline**, you will simulate the production usage of the model in a batch scenario. You will also implement data quality and stability checks, score model output, validate model quality, and use quality checks to make informed decisions on model retraining.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-b67ff9cdf73961d7714c2f104db2807f60da9d6b%2F202310_module5_fin.021-min.png?alt=media)

And now, to practice!


# 5.2. Train and evaluate an ML model \[OPTIONAL CODE PRACTICE]

A code example walkthrough of preparing the data, training, evaluating, and saving an ML model using the Evidently Python library.

{% embed url="<https://youtu.be/wHBL9zFgA8U?si=7jzJdynEt6tzAKLN>" %}

**Video 2**. [Train and evaluate an ML model \[OPTIONAL CODE PRACTICE\]](https://youtu.be/wHBL9zFgA8U?si=7jzJdynEt6tzAKLN), by Emeli Dral

In this video, we prepare the data, train, evaluate, and save an ML model that we will use later in this module to deploy an end-to-end pipeline for data and ML model quality checks.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module5/train_and_evaluate_model_practice.ipynb) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=0s) Introduction\
[00:39](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=39s) Imports\
[01:44](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=104s) Load and preview the raw data\
[05:02](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=302s) Feature engineering function\
[17:33](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=1053s) Split into train, reference and production\
[19:48](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=1188s) Transform raw data into pre-processed (and some debugging!)\
[23:00](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=1380s) Model training\
[27:28](https://www.youtube.com/watch?v=wHBL9zFgA8U\&t=1648s) Evaluate model quality


# 5.3. Test input data quality, stability and drift \[CODE PRACTICE]

A code example walkthrough of running test suites for data quality, data stability, and data drift on raw and pre-processed data.

{% embed url="<https://youtu.be/wZ0op3L9t2k?si=xQ-AoYiHp6zbv2Ml>" %}

**Video 3**. [Test input data quality, stability and drift \[CODE PRACTICE\]](https://youtu.be/wZ0op3L9t2k?si=xQ-AoYiHp6zbv2Ml), by Emeli Dral

In this video, we run test suites for data quality, data stability, and data drift on raw and pre-processed data. We also get the output as a Python dictionary to show how to integrate conditional checks in the prediction pipelines.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module5/data_quality_test_practice.ipynb) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=0s) Introduction\
[01:10](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=70s) Imports and data preparation\
[03:50](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=230s) Test data stability on raw data\
[06:40](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=400s) Run the test suite and explore the results\
[11:07](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=667s) Test data quality on raw data\
[12:34](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=754s) Test data drift on raw data\
[15:47](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=947s) Run tests and interpret data drift on pre-processed data\
[19:28](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=1168s) Whether to run tests on raw or pre-processed data\
[20:07](https://www.youtube.com/watch?v=wZ0op3L9t2k\&t=1207s) Get output as JSON or Python dictionary and create conditions


# 5.4. Test ML model outputs and quality \[CODE PRACTICE]

A code example walkthrough of testing the quality of the model outputs after generating predictions and getting the new labeled data.

{% embed url="<https://youtu.be/AcVhZeQvjSo?si=X9sleg-hV6kjbfJz>" %}

**Video 4**. [Test ML model outputs and quality \[CODE PRACTICE\]](https://youtu.be/AcVhZeQvjSo?si=X9sleg-hV6kjbfJz), by Emeli Dral

In this video, we test the quality of the model outputs after we generate the predictions. We also test the model quality after we get the new labeled data.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module5/model_quality_practice.ipynb) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=AcVhZeQvjSo\&t=0s) Introduction\
[00:44](https://www.youtube.com/watch?v=AcVhZeQvjSo\&t=44s) Imports\
[02:04](https://www.youtube.com/watch?v=AcVhZeQvjSo\&t=124s) Data loading and prep\
[06:10](https://www.youtube.com/watch?v=AcVhZeQvjSo\&t=370s) Load the model and get predictions\
[09:41](https://www.youtube.com/watch?v=AcVhZeQvjSo\&t=581s) Test model outputs and interpret the results\
[13:38](https://www.youtube.com/watch?v=AcVhZeQvjSo\&t=818s) Run tests for the next batch of data\
[15:20](https://www.youtube.com/watch?v=AcVhZeQvjSo\&t=920s) Test the model quality


# 5.5. Design a custom test suite with Evidently \[CODE PRACTICE]

A code example walkthrough of creating a custom Test Suite from individual tests and passing custom test conditions in the Evidently Python library.

{% embed url="<https://youtu.be/aE9ZlhaQ1fk?si=RLaXzZ3Jmi4nrPrn>" %}

**Video 5**. [Design a custom test suite with Evidently \[CODE PRACTICE\]](https://youtu.be/aE9ZlhaQ1fk?si=RLaXzZ3Jmi4nrPrn), by Emeli Dral

In this video, we show how to create a custom Test Suite from individual tests and pass custom test conditions.

**Want to go straight to code?** Here is the [example notebook](https://github.com/evidentlyai/ml_observability_course/blob/main/module5/custome_suite_practice.ipynb) to follow along.

**Link to docs:** [custom test suite](https://docs.evidentlyai.com/user-guide/tests-and-reports/custom-test-suite)

**Outline:**\
[00:00](https://www.youtube.com/watch?v=aE9ZlhaQ1fk\&t=0s) Introduction\
[00:50](https://www.youtube.com/watch?v=aE9ZlhaQ1fk\&t=50s) Imports, data loading and prep\
[03:05](https://www.youtube.com/watch?v=aE9ZlhaQ1fk\&t=185s) Load model and get predictions\
[04:26](https://www.youtube.com/watch?v=aE9ZlhaQ1fk\&t=266s) Create a custom test suite\
[09:06](https://www.youtube.com/watch?v=aE9ZlhaQ1fk\&t=546s) View the test results\
[10:04](https://www.youtube.com/watch?v=aE9ZlhaQ1fk\&t=604s) Run the test suite without reference data\
[12:59](https://www.youtube.com/watch?v=aE9ZlhaQ1fk\&t=779s) Recap and what's next


# 5.6. Run data drift and model quality checks in an Airflow pipeline \[OPTIONAL CODE PRACTICE]

A code example walkthrough of automating data and model quality checks implemented with the Evidently Python library using Airflow.

{% embed url="<https://youtu.be/YHO7k3T_fZA?si=9ePxgZso8mA4CsFR>" %}

**Video 6**. [Run data drift and model quality checks in an Airflow pipeline \[OPTIONAL CODE PRACTICE\]](https://youtu.be/YHO7k3T_fZA?si=9ePxgZso8mA4CsFR), by Emeli Dral

In this video, we show how to automate the data or model quality checks implemented with the Evidently Python library using Airflow.

**Want to go straight to code?** Here is the [code example](https://github.com/evidentlyai/ml_observability_course/tree/main/module5/airflow_conditional_checks) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=0s) Introduction\
[01:09](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=69s) Install Airflow\
[02:47](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=167s) Install dependencies\
[05:20](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=320s) Rebuild the container and access Airflow UI\
[07:06](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=426s) Start creating the DAG\
[10:00](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=600s) Specify DAG parameters\
[12:48](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=768s) Add functions and implement DAG\
[17:55](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=1075s) Implement load data function\
[19:21](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=1161s) Implement drift analysis function\
[21:32](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=1292s) Implement create report function\
[23:14](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=1394s) View DAG in Airflow\
[26:06](https://www.youtube.com/watch?v=YHO7k3T_fZA\&t=1566s) Execute a DAG and view the drift report


# 5.7. Run data drift and model quality checks in a Prefect pipeline \[OPTIONAL CODE PRACTICE]

A code example walkthrough of automating data and model quality checks implemented with the Evidently Python library using Prefect.

{% embed url="<https://youtu.be/ltmxxGV7Syg?si=efDXFxZNHndSidVI>" %}

**Video 7**. [Run data drift and model quality checks in a Prefect pipeline \[OPTIONAL CODE PRACTICE\]](https://youtu.be/ltmxxGV7Syg?si=efDXFxZNHndSidVI), by Emeli Dral

In this video, we show how to automate the data or model quality checks implemented with the Evidently Python library using Prefect.

**Want to go straight to code?** Here is the [code example](https://github.com/evidentlyai/ml_observability_course/tree/main/module5/prefect_sequential_checks) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=0s) Introduction\
[00:55](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=55s) Install the libraries\
[01:53](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=113s) Start creating a Prefect flow\
[03:58](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=238s) Create the flow structure\
[05:33](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=333s) Implement functions to load data and run test suites\
[11:12](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=672s) Run the Python script\
[13:09](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=789s) Implement the Prefect tasks and flow\
[15:36](https://www.youtube.com/watch?v=ltmxxGV7Syg\&t=936s) Run the flow and view it in Prefect UI


# 5.8. Log data drift test results to MLflow \[CODE PRACTICE]

A code example walkthrough of logging the results of data drift tests implemented with Evidently to Mlflow and viewing the results in the MLflow interface.

{% embed url="<https://youtu.be/gluRb9TbWSE?si=2Bw77DLhy_AKS-gz>" %}

**Video 8**. [Log data drift test results to MLflow \[CODE PRACTICE\]](https://youtu.be/gluRb9TbWSE?si=2Bw77DLhy_AKS-gz), by Emeli Dral

In this video, we show how to log the results of data drift tests implemented with Evidently to Mlflow and view the results in the MLflow interface.

**Want to go straight to code?** Here is the [code example](https://github.com/evidentlyai/ml_observability_course/tree/main/module5/mlflow_logging) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=0s) Introduction\
[00:50](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=50s) Start creating an MLflow script\
[01:48](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=108s) Set up the database and preview the MLflow UI\
[03:10](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=190s) Writing the Python script: overview and imports\
[04:24](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=264s) Writing the Python script: set experiments\
[08:03](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=483s) Writing the Python script: log parameters and reports\
[10:15](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=615s) Run the script and view the results\
[13:30](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=810s) Run a new experiment to log HTML reports\
[16:27](https://www.youtube.com/watch?v=gluRb9TbWSE\&t=987s) Next module focus

## Enjoyed the content?

Star Evidently on GitHub to contribute back! This helps us create free, open-source tools and content for the community.

⭐️ [Star](https://github.com/evidentlyai/evidently) on GitHub!


# Module 6: Deploying an ML monitoring dashboard

This module shows an end-to-end code example of designing and deploying an ML monitoring dashboard for batch and near real-time ML monitoring architectures.

As you deploy multiple models in production, you often want a live monitoring dashboard showing how all your ML models perform over time. The dashboard helps visualize the performance, detect issues, and debug them.

In this module, we will show an end-to-end example of designing and deploying an ML monitoring dashboard. We will cover both batch and near real-time model monitoring architectures. You will work with tools like Evidently and Grafana.

This is a code-focused module that includes hosting a local ML monitoring dashboard.


# 6.1. How to deploy a live ML monitoring dashboard

ML monitoring architectures recap and how to deploy a live ML monitoring dashboard.

{% embed url="<https://youtu.be/AZxp7f5IahU?si=JE-rf1g6iHoHSipU>" %}

**Video 1**. [How to deploy a live ML monitoring dashboard](https://youtu.be/AZxp7f5IahU?si=JE-rf1g6iHoHSipU), by Emeli Dral

## Why build a monitoring dashboard?

Tests and reports are great for running structured checks and exploring and debugging your data and ML models. However, when you run tests or build an ad-hoc report, you evaluate only a **specific batch of data**. This makes your monitoring **“static”**: you can get alerts but have no visibility into metric evolution and trends.

That is where the ML monitoring dashboard comes into play, as it:

* Tracks metrics over time,
* Aids in trend analysis and provides insights,
* Provides a shared UI to enhance visibility for stakeholders, e.g., data scientists, model users, product managers, business stakeholders, etc.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-186e62641cc6cb3a82c5e735086dda4635cb615c%2F202310_module6.004-min.png?alt=media)

## ML monitoring architectures recap

Before proceeding to the code practice of building a monitoring dashboard, let’s quickly recap ML monitoring architectures and how they differ.

ML monitoring starts with **logging**. To set up an effective ML monitoring backend for your service, you need to log input data, model outputs, and labeled data (if available) – they can be recorded in a **prediction store**.

You can send logged data to an **ML monitoring service** or run **monitoring jobs** to calculate monitoring metrics. Both methods allow recording calculated metrics in a **metric store**, which you can use as a data source to build a **monitoring dashboard**.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-4b3da728896c418b39905d1f171ce9fc53a470ea%2F202310_module6.006-min.png?alt=media)

## Code practice overview

We will show how to design and deploy an ML monitoring dashboard for batch and near real-time model monitoring architectures.

**For batch monitoring**, we will use Evidently open-source ML monitoring architecture. We will:

* Create **snapshots** that are individual Reports or Test Suites computed for a specific period (e.g., hour, day, week) in a rich JSON format.
* Log them to a file system and run an ML monitoring service that reads data from snapshots.
* Design dashboard panels and visualize **data and model metrics over time**. You can easily switch between time-series data and individual Reports/Test Suites for further analysis and debugging.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-23fe6fadd94737d61247c46a526d17607dbb9311%2F202310_module6.010-min.png?alt=media)

**For near real-time monitoring**, Evidently provides a **collector service**:

* It collects incoming data and computes Reports / Test Suites continuously.
* You can configure the service to calculate snapshots over various time intervals (as opposed to writing batch jobs on your own – you can define the frequency of batches using a config file).
* The service outputs the same snapshots as in the batch monitoring scheme: the front-end is the same.

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-84e588aa36fd755aa03635ced9aa936cde05bd7e%2F202310_module6.014-min.png?alt=media)

We will also cover an **alternative architecture using Grafana**, a popular open-source visualization tool. We will walk you through the Evidently and Grafana integration example – you can parse data computed by Evidently, store it in a chosen database, and visualize it in Grafana. In this case, Evidently serves as an evaluation layer.

This can be a good alternative if you already use Grafana for monitoring and alerting for other services.

And now, to practice!


# 6.2. ML model monitoring dashboard with Evidently. Batch architecture \[CODE PRACTICE]

A code example walkthrough of creating a live ML monitoring dashboard for batch architecture using Evidently.

{% embed url="<https://youtu.be/u4Mcu0hXfMA?si=uNdmwVdKcAWP410K>" %}

**Video 2**. [ML model monitoring dashboard with Evidently. Batch architecture \[CODE PRACTICE\]](https://youtu.be/u4Mcu0hXfMA?si=uNdmwVdKcAWP410K), by Emeli Dral

In this video, we create a script to generate Reports and Test Suites for several batches of data and design different panels to display on a live ML monitoring dashboard.

**Want to go straight to code?** Here is the [code example](https://github.com/evidentlyai/ml_observability_course/blob/main/module6/batch_monitoring_dashboard.py) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=0s) Introduction\
[00:47](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=47s) Overview of the script\
[02:27](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=147s) Imports\
[03:28](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=208s) Create global variables\
[05:41](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=341s) Load the data\
[07:19](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=439s) Implement the function to generate Reports and Test Suites\
[11:32](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=692s) Create a Project\
[13:53](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=833s) Add a counter panel (dashboard title)\
[15:10](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=910s) How to filter which data to display\
[16:31](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=991s) Add a line plot panel (target drift)\
[19:55](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1195s) Add a bar plot panel (dataset drift)\
[21:23](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1283s) Add test suite panels\
[22:47](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1367s) Implement the function to generate the Dashboard\
[23:46](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1426s) Live script debugging\
[27:11](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1631s) Run the script and monitoring service\
[28:36](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1716s) View and explore the dashboard in the browser\
[30:21](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1821s) View the individual Reports and Test Suites\
[31:38](https://www.youtube.com/watch?v=u4Mcu0hXfMA\&t=1898s) Recap and next steps


# 6.3. ML model monitoring dashboard with Evidently. Online architecture \[CODE PRACTICE]

A code example walkthrough of creating a live ML monitoring dashboard for online architecture using Evidently.

{% embed url="<https://youtu.be/2hTRXEOJF8k?si=CCVRwxiWWyZGmZF7>" %}

**Video 3**. [ML model monitoring dashboard with Evidently. Online architecture \[CODE PRACTICE\]](https://youtu.be/2hTRXEOJF8k?si=CCVRwxiWWyZGmZF7), by Emeli Dral

In this video, we create a live ML monitoring dashboard for an ML model deployed as a service. We imitate sending the live data directly from the machine learning service to the ML monitoring service and update the dashboard in near real-time.

**Want to go straight to code?** Here is the [code example](https://github.com/evidentlyai/ml_observability_course/blob/main/module6/online_monitoring_dashboard.py) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=0s) Introduction\
[00:30](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=30s) Script overview and imports\
[01:59](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=119s) Define Collector, Workspace, and Project variables\
[03:31](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=211s) Load data and create mini-batches to simulate production usage\
[04:39](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=279s) Implement the function to generate Test Suites\
[06:38](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=398s) Create the Workspace, Project and add Dashboard panels\
[09:25](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=565s) Set up and configure the Collector service\
[13:00](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=780s) Simulate sending data to the Collector\
[15:48](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=948s) Implement the main function, run and debug the script\
[18:32](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=1112s) Run the Collector and view the online Dashboard updates\
[20:46](https://www.youtube.com/watch?v=2hTRXEOJF8k\&t=1246s) Recap and next steps


# 6.4. ML monitoring with Evidently and Grafana \[OPTIONAL CODE PRACTICE]

A code example walkthrough of adding ML monitoring metrics to Grafana using Evidently as an evaluation layer.

{% embed url="<https://youtu.be/S4zFqbLhAp8?si=BEfLZteDmj94XPpD>" %}

**Video 4**. [ML monitoring with Evidently and Grafana \[OPTIONAL CODE PRACTICE\]](https://youtu.be/S4zFqbLhAp8?si=BEfLZteDmj94XPpD), by Emeli Dral

In this video, we show how to add ML monitoring metrics to the Grafana dashboard using Evidently as an evaluation layer and storing the metrics in a Postgres database.

**Want to go straight to code?** Here is the [code example](https://github.com/evidentlyai/ml_observability_course/tree/main/module6/grafana_monitoring_dashboard) to follow along.

**Outline:**\
[00:00](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=0s) Introduction\
[00:25](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=25s) Install Grafana and set up the Postgres database\
[03:18](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=198s) Script overview, imports, and logging\
[04:50](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=290s) Create a table statement\
[06:30](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=390s) Load data and simulate production usage\
[07:30](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=450s) Connect to a database and create a table\
[10:32](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=632s) Calculate metrics and insert them into Postgres\
[16:00](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=960s) Write a function to compute metrics in batches\
[18:58](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=1138s) Run services, execute and debug the script\
[23:21](https://www.youtube.com/watch?v=S4zFqbLhAp8\&t=1401s) Create dashboards in Grafana


# 6.5. Connecting the dots: full-stack ML observability

A brief summary of the Open-source ML observability course learnings.

{% embed url="<https://youtu.be/lK3OE473W0I?si=sPBZF_7wtUYgtRVO>" %}

**Video 5**. [Connecting the dots: full-stack ML observability](https://youtu.be/lK3OE473W0I?si=sPBZF_7wtUYgtRVO), by Emeli Dral

## What we covered in the course

This is the final lesson of the Open-source ML observability course. Let’s recap what we’ve learned during the course!

**ML monitoring metrics** We covered what metrics to use to assess [data quality](https://learn.evidentlyai.com/ml-observability-course/module-2-ml-monitoring-metrics/data-quality-in-ml), [model quality](https://learn.evidentlyai.com/ml-observability-course/module-2-ml-monitoring-metrics/ml-quality-metrics-classification-regression-ranking), and [data drift](https://learn.evidentlyai.com/ml-observability-course/module-2-ml-monitoring-metrics/data-prediction-drift-in-ml). We also discussed how to implement [custom metrics](https://learn.evidentlyai.com/ml-observability-course/module-4-designing-effective-ml-monitoring/custom-metrics-ml-monitoring) for specific use cases. For example, you can integrate custom metrics related to business KPIs and specific aspects of model quality into ML monitoring.

**ML monitoring design** We covered different aspects of ML monitoring design, including how to select and use [reference datasets](https://learn.evidentlyai.com/ml-observability-course/module-4-designing-effective-ml-monitoring/how-to-choose-reference-dataset-ml-monitoring). We also discussed the connection between [model retraining](https://learn.evidentlyai.com/ml-observability-course/module-4-designing-effective-ml-monitoring/when-to-retrain-ml-models) cadence and ML monitoring.

**ML monitoring architectures** We explored different [ML monitoring architectures](https://learn.evidentlyai.com/ml-observability-course/module-4-designing-effective-ml-monitoring/choosing-ml-monitoring-deployment-architecture), from ad hoc reports and test suites to batch and real-time ML monitoring, and learned how to implement them in practice in [Module 5](https://learn.evidentlyai.com/ml-observability-course/module-5-ml-pipelines-validation-and-testing) and [Module 6](https://learn.evidentlyai.com/ml-observability-course/module-6-deploying-an-ml-monitoring-dashboard).

**ML monitoring for unstructured data** We also touched on how to build a monitoring system for [text data](https://learn.evidentlyai.com/ml-observability-course/module-3-ml-monitoring-for-unstructured-data/monitoring-data-drift-on-raw-text-data) and [embeddings](https://learn.evidentlyai.com/ml-observability-course/module-3-ml-monitoring-for-unstructured-data/monitoring-embeddings-drift).

![](https://685625387-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FrV3hQYUjmKLt0fg4k4aS%2Fuploads%2Fgit-blob-b2e2435aeb01a56931d3e3b189cdf415e608e98a%2F202310_module6.022-min.png?alt=media)

## Summing up

**Start small and expand.** Ad hoc reports are a good starting point for ML monitoring that is easy to implement. It is useful for initial learning about data and model quality before establishing a comprehensive monitoring system. Don’t hesitate to start small!

As you progress and deploy multiple models in production, or if you work with mission-critical use cases, you’d need a more extensive setup.

**Jobs to be done** to implement full-stack production ML observability:

* **Immediate monitoring flow** helps to detect issues and to alert during model inference. If you have a production-critical service, it is essential to implement it.
* **Delayed monitoring flow** allows you to evaluate model quality when you get the labels (as ground truth is often not available immediately!).
* **Model evaluation flow** is needed to test model quality at updates and retraining.

**Observability components** to keep in mind when building ML monitoring:

* **Logging layer**. If you have a production service, implementing logging is a must to capture model inferences and collect performance metrics.
* **Alerting layer** allows you to monitor metrics and get notifications when things go wrong.
* **Dashboarding and analytics** help to visualize the performance, quickly detect root cause issues, and define actions for debugging and retraining.

## Enjoyed the course?

⭐️ [**Star Evidently on GitHub**](https://github.com/evidentlyai/evidently) to contribute back! This helps us create free, open-source tools and content for the community.

📌 [**Share your feedback**](https://db984wnn7iq.typeform.com/to/IKDoGgKx) so we can make this course better.

💻 [**Join our Discord community**](https://discord.com/invite/xZjKRaNp8b) for more discussions and materials on ML monitoring and observability.


