Skip to main content
Back to Blog
Tutorial
2026-09-28

Testing Apache Airflow DAGs with pytest: Integrity, Unit, and Integration Tests

Airflow DAG testing guide for pytest users: validate imports, task logic, dag.test runs, fixtures, mocks, Docker checks, CI gates, and data contracts.

Testing Apache Airflow DAGs with pytest: Integrity, Unit, and Integration Tests

Airflow DAG testing is the practice of proving that workflow files import cleanly, define the graph you expect, run task logic with controlled inputs, and survive a realistic Airflow runtime before they reach the scheduler. In 2026, Apache Airflow is actively maintained, Airflow 3.x is the current major line, and the latest release I verified is 3.3.2 from September 17, 2026. The most important Airflow 3 testing change is that DAG authors should import authoring primitives from airflow.sdk, while many legacy Airflow 2 examples still import directly from older internal modules.

Use pytest as the control plane for three layers: DAG integrity tests with DagBag, unit tests for pure Python transformation code and TaskFlow callables, and integration tests that execute a small DAG run through dag.test() or airflow dags test. The payoff is not just a green CI job. It is a reviewable contract for schedules, task IDs, retries, pools, variables, connections, side effects, and data quality handoffs.

This guide assumes you are a QA or test automation engineer maintaining DAGs with help from Claude Code, Cursor, Copilot, or another AI coding agent. The agent can write a lot of boilerplate quickly, but it will also copy old Airflow 2 imports, invent provider paths, or turn every task into a slow integration test unless you give it a layered test strategy.

If your DAGs orchestrate data checks rather than only application jobs, pair the workflow tests here with metamorphic testing for data pipelines and Great Expectations data quality testing. Those checks validate the data contract after Airflow proves the schedule, graph, and runtime wiring.

The Current Airflow Testing Surface

Airflow 3 did not discontinue DAG testing. It did change the safest authoring surface. The official Airflow 3 public interface documentation says DAG authors should use the airflow.sdk namespace for objects such as DAG, @dag, @task, Variable, Connection, and get_current_context. The Airflow 3.0 release notes also call out provider import movement, including standard operators under airflow.providers.standard.

DagBag still exists and is documented as the collection that parses DAGs from a folder tree, but Airflow labels it as an internal loading mechanism rather than the preferred authoring API. That nuance matters: use DagBag in tests to catch import failures and duplicate IDs, but do not teach authors to build DAGs through DagBag.

Test layerWhat it provesFast signalFailure that it catches
Static lint and importDAG files are syntactically valid and importableSecondsRemoved Airflow 2 import, missing provider package, accidental network call at import time
DAG integrityIDs, schedules, owners, tags, and dependencies match policySeconds to a minuteDuplicate task ID, unpaused catchup DAG, unexpected schedule
Unit testsBusiness logic works without a schedulerSecondsBad date parsing, null handling, SQL builder bug, malformed payload
Local DAG executionAirflow runtime can execute a tiny runMinutesXCom mismatch, templating error, connection lookup error
Container integrationThe same image and providers work in CIMinutes to tens of minutesProvider not installed in image, DB migration issue, runtime-only dependency gap

What people get wrong: they start by running a full DAG as the first test. That makes the suite slow, hides the cause of failures, and encourages broad mocking of Airflow internals. Start with import and graph contracts, unit test the code that actually transforms data, then run one or two execution paths that represent release risk.

Project Layout That Makes DAGs Testable

A testable Airflow repository separates orchestration from business logic. The DAG file should be boring: declare schedule, tasks, dependencies, and small TaskFlow wrappers. The transformation code should live in normal Python modules that pytest can import without creating an Airflow metadata database.

PathResponsibilityTest style
dags/orders_daily.pyDAG definition, TaskFlow wrappers, dependenciesDagBag import and graph assertions
src/orders/normalize.pyPure data validation and normalizationPlain pytest unit tests
tests/dag_integrity/DAG policy checks across all DAGsShared DagBag fixture
tests/unit/Fast tests for task logicParametrized pytest tests
tests/integration/dag.test() and CLI smoke checksMarked as integration
docker-compose.ymlAirflow services for local and CI smokeSmall runtime verification
[project]
name = "airflow-dag-tests"
version = "0.1.0"
requires-python = ">=3.11,<3.13"
dependencies = [
  "apache-airflow==3.3.2",
  "apache-airflow-providers-standard==1.19.0",
  "pendulum>=3.0.0",
]

[dependency-groups]
dev = [
  "pytest>=8.0.0",
  "ruff>=0.13.0",
]

[tool.pytest.ini_options]
testpaths = ["tests"]
markers = [
  "integration: requires an initialized Airflow runtime",
  "dag_integrity: parses DAG files and validates graph policy",
]
addopts = "-ra --strict-markers"

The provider pin above is illustrative. Verify provider versions against your lockfile, because Airflow providers release independently. The important pattern is that the same provider set used by the scheduler should be installed in CI. A DAG that imports locally but fails inside the production image is usually a packaging problem, not a pytest problem.

A Small DAG With Testable Task Logic

The cleanest DAG test is the one that does not need Airflow at all. Put domain behavior in a normal function, wrap it with @task, and test the function directly. The DAG below uses Airflow 3 authoring imports, a deterministic start date, and no top-level API calls.

# dags/orders_daily.py
from __future__ import annotations

import pendulum
from airflow.sdk import dag, task


def normalize_order(raw: dict[str, object]) -> dict[str, object]:
    order_id = str(raw.get("order_id", "")).strip()
    country = str(raw.get("country", "")).strip().upper()
    cents = int(raw.get("amount_cents", 0))
    if not order_id:
        raise ValueError("order_id is required")
    if cents < 0:
        raise ValueError("amount_cents must be non-negative")
    return {
        "order_id": order_id,
        "country": country,
        "amount_cents": cents,
    }


@dag(
    dag_id="orders_daily",
    schedule="@daily",
    start_date=pendulum.datetime(2026, 1, 1, tz="UTC"),
    catchup=False,
    tags=["orders", "qa-owned"],
)
def orders_daily():
    @task
    def extract() -> dict[str, object]:
        return {"order_id": "A-100", "country": "us", "amount_cents": 1299}

    @task
    def normalize(raw: dict[str, object]) -> dict[str, object]:
        return normalize_order(raw)

    @task
    def publish(order: dict[str, object]) -> str:
        return str(order["order_id"])

    publish(normalize(extract()))


orders_daily()

This shape gives an AI coding agent a safe target: change normalize_order when the data rule changes, and change the DAG only when orchestration changes. If the agent edits both at once, reviewers can ask which layer failed and which test proves the fix.

DAG Integrity Tests With DagBag

DagBag tests are the quickest way to catch broken imports before the scheduler does. They should answer concrete questions: did every DAG parse, did the expected DAG load, did the graph contain the expected task IDs, and did policy constraints hold? They should not execute the whole pipeline.

# tests/conftest.py
from __future__ import annotations

from pathlib import Path

import pytest
from airflow.models.dagbag import DagBag


PROJECT_ROOT = Path(__file__).resolve().parents[1]
DAGS_DIR = PROJECT_ROOT / "dags"


@pytest.fixture(scope="session")
def dagbag() -> DagBag:
    return DagBag(dag_folder=str(DAGS_DIR), include_examples=False)
# tests/dag_integrity/test_orders_daily.py
from __future__ import annotations


def test_dagbag_has_no_import_errors(dagbag):
    assert dagbag.import_errors == {}


def test_orders_daily_graph_contract(dagbag):
    dag = dagbag.get_dag("orders_daily")

    assert dag is not None
    assert dag.schedule == "@daily"
    assert dag.catchup is False
    assert {"orders", "qa-owned"}.issubset(set(dag.tags))
    assert sorted(task.task_id for task in dag.tasks) == [
        "extract",
        "normalize",
        "publish",
    ]
    assert dag.get_task("publish").upstream_task_ids == {"normalize"}
    assert dag.get_task("normalize").upstream_task_ids == {"extract"}

Two details are deliberate. First, the import error assertion compares to an empty dictionary, so the failure prints the exact files and tracebacks. Second, the dependency assertions check both presence and direction. A weak test that only checks len(dag.tasks) == 3 can pass after a dependency is accidentally removed.

Unit Testing TaskFlow Functions Without A Scheduler

TaskFlow makes it easy to blur the line between Python functions and Airflow tasks. Keep one pure function behind every nontrivial task, and unit test that function. You do not need a scheduler to prove basic rules.

# tests/unit/test_normalize_order.py
from __future__ import annotations

import pytest

from dags.orders_daily import normalize_order


@pytest.mark.parametrize(
    ("raw", "expected"),
    [
        (
            {"order_id": " A-100 ", "country": "us", "amount_cents": "1299"},
            {"order_id": "A-100", "country": "US", "amount_cents": 1299},
        ),
        (
            {"order_id": "B-200", "country": "in", "amount_cents": 0},
            {"order_id": "B-200", "country": "IN", "amount_cents": 0},
        ),
    ],
)
def test_normalize_order_returns_canonical_payload(raw, expected):
    assert normalize_order(raw) == expected


def test_normalize_order_rejects_missing_order_id():
    with pytest.raises(ValueError, match="order_id is required"):
        normalize_order({"country": "US", "amount_cents": 100})


def test_normalize_order_rejects_negative_amount():
    with pytest.raises(ValueError, match="amount_cents must be non-negative"):
        normalize_order({"order_id": "A-100", "country": "US", "amount_cents": -1})

These are the tests you want an agent to run after every small edit. They are cheap, deterministic, and meaningful. The exception checks are anchored to the behavior, not just the exception type, which keeps accidental broad exceptions from looking correct.

Testing Variables, Connections, And Runtime Context

Airflow variables and connections create most local testing surprises. The official connection docs confirm that environment-backed connections use AIRFLOW_CONN_{CONN_ID} in uppercase and can contain either a URI or JSON value. Airflow variables can also be supplied as environment variables with AIRFLOW_VAR_{KEY}. Environment-backed values may not appear in the UI, which is a feature for secret injection, not a sign that the test setup failed.

Runtime dependencyUnit test approachIntegration test approachCommon mistake
Connection URIInject a function argument or set AIRFLOW_CONN_...Use a test secret or local service URICreating real production connections in CI
VariablePass config as data or set AIRFLOW_VAR_...Seed a value before dag.test()Reading variables at module import time
Current contextPass dates explicitly to pure functionsUse get_current_context() inside a taskCalling context APIs outside task execution
External APIFake the client and assert request payloadHit a sandbox endpoint only in marked testsMocking so deeply that no payload is verified
# tests/unit/test_runtime_config.py
from __future__ import annotations

from airflow.sdk import Variable


def bucket_name() -> str:
    return Variable.get("raw_orders_bucket")


def test_bucket_name_comes_from_airflow_variable(monkeypatch):
    monkeypatch.setenv("AIRFLOW_VAR_RAW_ORDERS_BUCKET", "qa-orders-landing")

    assert bucket_name() == "qa-orders-landing"

The important testing rule is to avoid top-level runtime lookups in DAG files. If a DAG reads a variable during import, a missing value breaks DagBag parsing and the scheduler may fail before a task has a chance to run. Read variables inside tasks or pass configuration through stable deployment mechanisms.

Running A DAG Locally With dag.test

The official debugging docs describe dag.test() as a way to run a DAG in a single serialized Python process. It can run locally without an executor by default, and it accepts options such as an execution date and use_executor. The CLI reference also documents airflow dags test, including --dagfile-path, --conf, --mark-success-pattern, --save-dagrun, --show-dagrun, and --use-executor.

# scripts/run_orders_daily.py
from __future__ import annotations

import pendulum

from dags.orders_daily import orders_daily


dag = orders_daily()

if __name__ == "__main__":
    dag.test(execution_date=pendulum.datetime(2026, 9, 28, tz="UTC"))
airflow dags test orders_daily 2026-09-28 --dagfile-path dags/orders_daily.py --mark-success-pattern '^wait_for_'

dag.test() is excellent for debugging because it fails close to the Python code. The CLI is better when you want a CI check that resembles the operator workflow and uses the initialized Airflow home. Use --mark-success-pattern only for tasks that are intentionally out of scope, such as a sensor that waits on a production-only dependency. Do not mark a broken transform as successful just to keep a smoke test green.

Container Integration With Docker Compose

Unit tests can pass while the Airflow image is missing a provider package. A container smoke test catches that gap. Keep it narrow: initialize the metadata database, parse DAGs, then test one tiny DAG run. The goal is not to recreate production. It is to prove the image, providers, environment, and DAG folder agree.

CheckCommand shapeGood signalBad signal
Airflow versionairflow versionImage uses the pinned releaseLocal pip version differs from image
DB migrationairflow db migrateMetadata schema initializesConstraint or dependency conflict
DAG listairflow dags listDAG imports in containerMissing provider, import-time variable read
DAG executionairflow dags testRuntime templating and task execution workConnection, XCom, or provider mismatch
# docker-compose.test.yml
services:
  airflow-test:
    image: apache/airflow:3.3.2
    environment:
      AIRFLOW__CORE__LOAD_EXAMPLES: "false"
      AIRFLOW__CORE__EXECUTOR: LocalExecutor
      AIRFLOW__DATABASE__SQL_ALCHEMY_CONN: sqlite:////tmp/airflow-ci.db
      AIRFLOW_VAR_RAW_ORDERS_BUCKET: qa-orders-landing
      AIRFLOW_CONN_WAREHOUSE: sqlite:////tmp/warehouse.db
    volumes:
      - ./dags:/opt/airflow/dags:ro
      - ./src:/opt/airflow/src:ro
    command: >
      bash -c "airflow db migrate &&
      airflow dags list &&
      airflow dags test orders_daily 2026-09-28 --dagfile-path /opt/airflow/dags/orders_daily.py"

If this fails with an import error that pytest did not catch, compare the Python path and installed packages between local tests and the container. If it fails only at task execution, look for context, secrets, templates, XCom serialization, or provider-specific runtime dependencies.

CI Gates For Pull Requests And Nightly Runs

Pull request checks should be fast enough that developers trust them. Nightly checks can afford Docker and broader DAG execution. GitHub Actions current majors include actions/checkout@v7, actions/setup-python@v7, and actions/upload-artifact@v7; pin your own dependency versions with a lockfile or constraints.

name: airflow-dag-tests

on:
  pull_request:
  push:
    branches: [main]
  schedule:
    - cron: "17 2 * * *"

jobs:
  pytest:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: actions/setup-python@v7
        with:
          python-version: "3.11"
      - name: Install dependencies
        run: |
          python -m pip install --upgrade pip
          python -m pip install -e ".[dev]"
      - name: Static checks
        run: ruff check dags src tests
      - name: Fast pytest
        run: pytest -m "not integration" --maxfail=3

  airflow-runtime-smoke:
    runs-on: ubuntu-latest
    if: github.event_name != 'pull_request'
    steps:
      - uses: actions/checkout@v7
      - name: Run container smoke
        run: docker compose -f docker-compose.test.yml up --abort-on-container-exit --exit-code-from airflow-test
      - uses: actions/upload-artifact@v7
        if: always()
        with:
          name: airflow-runtime-${{ github.run_id }}
          path: logs
          if-no-files-found: ignore

The split is intentional. Pull requests run pytest -m "not integration", while scheduled or main-branch checks run the container smoke. If your repository has hundreds of DAGs, add changed-file selection later, but keep at least one full parse run on a schedule so hidden import drift does not accumulate.

Diagnosing A Realistic Failure Mode

Imagine CI reports this failure:

E   ModuleNotFoundError: No module named 'airflow.operators.python'

Diagnosis in Airflow 3: many examples written for Airflow 2 imported PythonOperator from airflow.operators.python. Airflow 3 release notes say legacy imports are deprecated and standard operators should come from the standard provider. The fix is to install the provider used by production and update the import path, or better, move simple Python work to TaskFlow with @task.

# Airflow 3 standard provider import when an operator is still needed
from airflow.providers.standard.operators.python import PythonOperator

Another common failure is KeyError or AirflowNotFoundException during DagBag import because a DAG reads a variable at module import time. Move that lookup inside a task, or supply the variable through an environment variable in the test process. Import tests should be boring; runtime configuration belongs at runtime.

Where Agent-Generated Tests Need Review

AI coding agents are useful for expanding fixtures and table-driven tests, but they tend to over-mock Airflow. Ask the agent for tests that assert graph contracts and side effects, not just that a method was called. For example, a publish task test should verify the exact payload or target partition. A DAG integrity test should verify dependency direction. A retry policy test should verify retry delay and owner, not only that the DAG exists.

Ready-made QA skills install from qaskills.sh with the qaskills CLI, but the skill still needs repository-specific constraints: Airflow version, provider set, DAG folder, secrets strategy, CI runner, and which DAGs are allowed to hit external services.

The best guardrail is a checklist in the prompt: use Airflow 3 imports, avoid top-level network calls, keep unit tests independent of the scheduler, use pytest markers for integration, and run DagBag before container smoke. That turns the agent from a snippet generator into a test maintainer.

Frequently Asked Questions

Should I use DagBag in Airflow 3 tests if it is internal?

Yes, for integrity tests, with the right framing. Airflow 3 documentation describes DagBag as the loader that parses DAGs from folders, and notes that DAG authors should use airflow.sdk for authoring. That means DagBag is reasonable in a test harness that checks import errors and graph contracts. It should not become the pattern for writing DAGs. Keep authoring examples on @dag, @task, and other public airflow.sdk imports.

Is dag.test enough for production confidence?

dag.test() is a strong local execution check, but it is not a complete production simulation. By default it runs tasks locally in one process, which is exactly why it is useful for debugging. It does not prove your Kubernetes executor, Celery workers, external secrets backend, or production network policy. Use it for fast runtime confidence, then add a narrow container or environment smoke test for image, provider, metadata database, and secret wiring.

How many DAG integration tests should run on every pull request?

Run the smallest set that catches release-blocking mistakes without making developers wait. A good default is all import and DAG integrity tests on every pull request, pure unit tests for touched packages, and one or two integration tests only when they are deterministic and under a few minutes. Put broader airflow dags test and Docker Compose checks on main or nightly. Slow, flaky PR gates train teams to ignore red builds.

What is the safest way to mock Airflow connections and variables?

Prefer environment-backed values for tests that need Airflow lookup behavior. Use AIRFLOW_CONN_MY_CONN_ID for connections and AIRFLOW_VAR_MY_KEY for variables, with uppercase names. For pure unit tests, pass configuration as function arguments instead of calling Airflow APIs. Avoid writing production-like secrets into the metadata database during CI. Also avoid reading variables at DAG import time, because that turns configuration absence into a parsing failure.