This guide provides ideals for production-grade Python code, code that humans are likely to read, code that will likely be edited and/or re-executed at least once, and code that I will review in detail. Code whose correctness, reproducibility, and reliability matter should follow these guidelines.
This guide partly exists because Python is so flexible. There are enough ways of using Python to warrant a clean separation between “scripts that will be executed once, never reviewed by another person, or used again” from “services or ‘important’ tasks whose correctness matters and for which I might be called in to debug at midnight.”
Containers
Cont.1. Prefer immutable collections and data structures to mutable ones especially to represent “facts known at a particular instant”.
Mutable collections are great for accreting information over time. Though when possible, accrete that state in one scope so understanding “how did a collection get to be how it is” only requires local analysis and reasoning.
Generally we agree with P.10 of the C++ Core Guidelines and its reasoning.
Reason: It is easier to reason about constants than about variables. Something immutable cannot change unexpectedly. Sometimes immutability enables better optimization. You can’t have a data race on a constant.
Cont.2. Prefer NamedTuples and (frozen) dataclasses for most “plain-old-data” containers over bare classes
Reason: typing.NamedTuples and frozen dataclasses.dataclasses signal to
experienced programmers that a “struct” is a “fact known at a particular point
in time.” They can be more easily and reliably passed across async and picked
across interprocess boundaries than many kinds of classes. Even
unfrozen/“normal” dataclasses signal to experienced developers that “this
container is mainly for holding data.” Bare/raw classes are better for state
machines or crafting a Pythonic API with instance-scoped encapsulated
implementation details. Ultimately an experienced developer can skim and fully
understand a NamedTuple or (frozen) dataclass more quickly than a raw/bare
class.
Example: A dataframe should be a class because it typically consists of a crafted “Pythonic” API over instance-scoped encapsulated/hidden “struct-of-arrays”-oriented numpy arrays or arrow record batches. A struct used to represent a validated collection of RPC request parameters that will be passed to a service might be better as a dataclass.
Cont.3. Use TypedDict for externally-enforced fixed-structure existing dictionaries.
Reason: typing.TypedDict was designed for adding type annotations/static
analyzability for existing untyped code bases without sweeping structural
changes.
PEP-589 that introduced typing.TypedDict
states the following.
This PEP proposes a type constructor
typing.TypedDictto support the use case where a dictionary object has a specific set of string keys, each with a value of a specific type…Dataclasses are a more recent alternative to solve this use case, but there is still a lot of existing code that was written before dataclasses became available, especially in large existing codebases where type hinting and checking has proven to be helpful. Unlike dictionary objects, dataclasses don’t directly support JSON serialization
typing.TypedDict is a more descriptive and less prescriptive kind of
documentation than a typing.NamedTuple or (frozen) dataclasses.dataclass.
typing.TypedDict is extremely useful for adding static analysis to “legacy”
applications where something outside the scope of the Python program creates the
dictionary. For example, consider a proprietary service layer where highly tuned
C++ written 30 years ago handles all schema enforcement and (de)serialization,
and the choice was made long ago to expose these to Python using dictionaries.
typing.TypedDict can be a useful code generation target because a system
outside the scope of the Python program reliably handles schema enforcement and
the decision to use Python dictionaries as the interface with the C++ layer is
outside the scope of the Python program.
typing.TypedDict can also be useful for adding type annotations to an existing
untyped code base that passes untyped dictionaries across function boundaries –
or worse, module boundaries. Consider a function in such a code base that
returns some dictionary where precise key names and value types can be inferred
from one function or module scope. Changing to a dataclass or NamedTuple may
prove difficult because one would need to change all call sites of the function
– a rather dangerous endeavor without strong static analysis and tests. One can
create a typing.TypedDict for the return type of the function and at least achieve
some type inference and static analysis around the call sites within one
commit/atomic change. Then after strengthening the test suite and type
annotations for the code base (e.g. passing mypy strict mode), one can more
confidently refactor to use dataclasses or NamedTuples.
Input-Output
IO.1. Validate untested/external data at runtime as close to the external boundary as possible
Parse HTTP payloads, queue messages, configuration files, environment variables,
plugin inputs, and third-party responses into validated domain values as near to
the external boundary as possible. Use pydantic or dataclass-based adapters,
or explicit parsing depending on schema complexity and project dependencies.
Static types typically do not validate runtime values. They complement, but do
not replace, boundary validation.
Reason: Detecting unmet expectations early maximizes the chances of good
error reporting with sufficient context for debugging. Avoid scenarios where a
program uses external data that doesn’t meet expectations for much of the
program and not knowing whether a KeyError or AttributeError is caused by
garbage data or a programmer bug.
The industry has largely settled on pydantic for strict parsing of
complex/nested data from outside the scope of the Python program, such as API
responses, user/client-provided configurations and forms, and queue messages
because it is fast and uses good idioms for avoiding possibly hundreds or
thousands of lines of parsing code. Pydantic, as long as one parses external
data as soon as possible, can help detect violations of contracts/expectations
can be reliably detected and handled.
The key here is that data that is generated within the Python program should
be able to use dataclasses and NamedTuples with static type checking to get
the often formidable “enforcement” afforded by well-tuned static analysis and
CI. There is generally no need to parse nested data structures entirely
generated and consumed within the same Python program because type checkers
detect expectation/contract violations at struct-creation time and
access-time.
Exceptions: Small scripts that choose to be single-file and zero-dependency
can parse hand-written, trusted, and controlled configuration files manually.
For a few key-value pairs, especially ones that are mostly strings and integers,
keep the parsing to a single function. pydantic makes more sense once the
cost of packaging is already being paid but it is typically not worth it if
bringing in pydantic creates the cost of packaging.
from dataclasses import dataclass
from pathlib import Path
@dataclass(frozen=True)
class _BackupRestoreValidationConfig: ...
def _parse_toml_config_file(config_path: Path) -> _BackupRestoreValidationConfig: ...
Control Flow
CF.1. Prefer context managers for scoped resource access and ‘computational contexts’ over try-finally
Reason: Context managers idiomatically enable setup-teardown patterns and are a reliable and simple way to encapsulate the process of releasing resources/tearing down execution context. This applies to many external resources, such as file handles, database connections and transactions, locks, and sockets/TCP connections.
Bad Example
import sqlite3
from dataclasses import dataclass
@dataclass(frozen=True)
class _DBConfig: # stand in for something akin to a connection pool
db_uri: str
def _get_username(user_id: int, config: _DBConfig) -> str | None:
conn = sqlite3.connect(config.db_uri)
cursor = conn.cursor()
try:
maybe_user_name = cursor.execute(
"SELECT username FROM users WHERE user_id=?",
(user_id,),
).fetchone()
return maybe_user_name[0] if maybe_user_name else None
finally:
cursor.close()
conn.close()
This is bad because there are situations where the connection and/or cursor will
not be reliably closed. For example, if cursor.close() fails, the connection
will leak. Moreover, failing to encapsulate the resource freeing in a context
manager forces more functions to also implement the clean-up. The try-finally
also complicates the swift closing/releasing of the connection.
Example
from contextlib import closing
import sqlite3
from dataclasses import dataclass
@dataclass(frozen=True)
class _DBConfig:
db_uri: str
def _get_username(user_id: int, config: _DBConfig) -> str | None:
with closing(sqlite3.connect(config.db_uri)) as conn:
maybe_user_name = conn.execute(
"SELECT username FROM users WHERE user_id=?",
(user_id,),
).fetchone()
return maybe_user_name[0] if maybe_user_name else None
A subtle difference here is that context managers help release the resources sooner, which can matter in certain demanding applications. While holding the connection for an extra ternary expression likely doesn’t matter, closing the connection early can matter with more demanding data processing.
Database client libraries, especially DBAPI-compatible ones, sometimes differ in
their interpretation of context-manager usage. Some, such as sqlite3 and
pyodbc, use the Connection.__exit__ method for transaction management and
roll back transactions on errors, while using the Connection.close method
(usually via with contextlib.closing(connect(....)) as conn) for closing the
connection. Others, such as psycopg3, use the Connection.__exit__ method to
roll back transactions on errors AND close the connection, making the invocation
of the Connection.close method via contextlib.closing unnecessary.
Connection pools, such as those provided by sqlalchemy have
Connection.__exit__ return the connection back to the connection pool. Pay
careful attention to how database libraries implement their context managers.
Typically Cursor.__exit__ closes the cursor/stops receiving rows from the
database and prevents further use, though this is underspecified by the dbapi2
spec, leaving implementations to decide the desired behavior of the cursor
context manager.
Further, note that for long-running production services, the better
recommendation is usually to use a database driver/pool whose lifecycle and
transaction semantics are clearly understood, and manage it at application
startup/shutdown rather than opening a new connection per request. One or a
couple of isolated connection-per-SELECT queries is fine for scripts.
Bad Example
import json
def read_config_file(file_name: str) -> dict[str, str]:
fob = open(file_name)
try:
return json.load(fob)
finally:
fob.close()
Example
import json
def read_config_file(file_name: str) -> dict[str, str]:
with open(file_name) as fob:
return json.load(fob)
The second block more concisely expresses and guarantees (under most, though not
all scenarios) that the file handle must be closed. Especially when the
‘computational context’ extends across more statements, context managers often
more reliably tear down the context/release the resource because the “setup”
and “teardown” logic is adjacent – in the class definition __enter__ and
__exit__ method or in a contextlib.contextmanager-decorated function –
rather than separated by potentially many lines of code. This pattern applies to
many kinds of computational context apart from just resource management.
Bad Example
import shutil
from logging import Logger
import time
def copy_directory_and_log_timing(src: str, dest: str, logger: Logger) -> None:
start = time.monotonic()
shutil.copytree(src, dest, dirs_exist_ok=True)
end = time.monotonic()
logger.info(
"Copied %s to %s",
src,
dest,
extra={"copy_duration_ms": int((end - start) * 1000)},
)
Example
import shutil
from logging import Logger
from collections.abc import Iterator
from contextlib import contextmanager
import time
@contextmanager
def log_timing(
message: str, logger: Logger, duration_ms_extra_key: str
) -> Iterator[None]:
start = time.monotonic()
yield
end = time.monotonic()
logger.info(
message,
extra={duration_ms_extra_key: int((end - start) * 1000)},
)
def copy_directory_and_log_timing(src: str, dest: str, logger: Logger) -> None:
with log_timing(
f"Copied {src} to {dest}", logger, duration_ms_extra_key="copy_duration_ms"
):
shutil.copytree(src, dest, dirs_exist_ok=True)
This change simplifies the process of instrumenting/logging timings of other operations.
With the contextlib.contextmanager decorator and its async sister
contextlib.asynccontextmanager, we almost never need to write explicit
__enter__ and __exit__ methods with their complex function signatures, as we
have a very simple pattern for separating the setup and teardown with a yield
statement.
Error Handling
Error states abound in production, especially for long-running processes. External services/dependencies can have downtime, performance regressions, and backwards-incompatible schema changes. Writing production-grade Python programs involves deliberately handling many failure modes.
- For errors arising from a programmer bug or invalid application state (e.g.
deadlock), there is no need to
try-exceptbecause such situations should crash the program in many cases, or at least fail loudly. - Some errors can be handled within the “responsibility scope” of the function or module. For example, we might retry ephemerally failing HTTP/service requests with backoff (and possibly jitter), especially if we know that the API being requested is particularly flaky.
- Some errors are better handled one or more layers up the call stack, in which case, we might use an errors-as-values pattern and propagate the stable/well-understood/controlled domain errors up the call stack via return values and let the caller (attempt to) recover or retry.
Do not continue when process-wide correctness is uncertain. Do not terminate healthy work unnecessarily when the failure is safely isolated.
In general, catch an exception only when the Python program can do one of the following:
- recover safely
- translate the exception into a stable domain error (especially as a value passed through the call stack via return values)
- add essential context while preserving the cause
- perform required cleanup not otherwise managed by a context manager
Err.1. Catch exceptions as specifically as possible
Reason: Robust programs consider at least one to three failure states/situations of each expression/procedure/function that is invoked. most-common failure cases. For each of the states and account for them
Bad Example
import json
from logging import Logger
def _read_val_from_config_file(file_name: str, logger: Logger) -> int | None:
try:
with open(file_name) as fob:
data = json.load(fob)
return int(data["some"]["key"]["that"]["may"]["not"]["exist"])
except Exception:
logger.exception("Something went wrong reading a value from %s", file_name)
return None
Most exceptions are not very exceptional. In the above cases, most file names are not real files that exist on the current machine; a given user cannot open most files on a system; most files are not JSON, most JSON files do not have the used keys. The “exceptional” cases are actually common and should have explicit error messages so a user can easily address issues. Production-grade, reliable scripts and especially long-running services do not have the luxury of only considering the “happy-path” of a program.
Example
import json
from logging import Logger
def _read_val_from_config_file(file_name: str, logger: Logger) -> int | None:
try:
with open(file_name) as fob:
data = json.load(fob)
except json.JSONDecodeError:
logger.exception("File %s is not valid JSON", file_name)
return None
except FileNotFoundError:
logger.error("File %s not found", file_name)
return None
try:
raw_extracted = data["some"]["key"]["that"]["may"]["not"]["exist"]
except (KeyError, AttributeError):
logger.exception("Failed to access data in JSON file %s", file_name)
return None
try:
parsed = int(raw_extracted)
except ValueError:
logger.error(
"Failed to parse value %s to integer in file %s", raw_extracted, file_name
)
return None
return parsed
Exceptions: Bare except Exception can be reasonable very close to the
entrypoint of a program to guarantee some sort of external behavior even in the
presence of unhandled exceptions. Web frameworks frequently implement this in
middleware to return a 500 HTTP status code and keep the server alive.
Err.2. Keep try-except blocks tight and surrounding as few failure-modes as possible
Reason try-except blocks that cover many statements complicate error
attribution and the creation of helpful error messages. In a
try-except-(finally)-block covering 5+ statements, many operations could
throw, for example, a ValueError, which complicates the creation of a
targeted, maximally useful error message.
Bad Example
In this example, consider a script that runs, for example, once per day and typically exits within 5 seconds.
# /// script
# requires-python = ">=3.11"
# dependencies = ["polars", "requests"]
# ///
from pathlib import Path
from datetime import date
import logging
import requests
import polars as pl
def etl_airline_data_to_parquet(
base_url: str,
output_dir: Path,
publish_date: date,
logger: logging.Logger,
) -> Path | None:
"""Download data from airline API to hive-style partitioned parquet for
a date that has not been processed yet, returning the newly written parquet
file if it was written, otherwise None if data was already processed/written"""
try:
output_pq_path = output_dir.joinpath(
f"publish_date={publish_date.isoformat()}", "data.parquet"
)
if output_pq_path.is_file():
return None
resp_json = requests.get(
base_url, data={"publish_date": publish_date.isoformat()}
).json()
output_pq_path.parent.mkdir(exist_ok=True, parents=True)
_schema = {"flight_id": pl.Int64, "passengers": pl.Int64, "revenue": pl.Int64}
pl.LazyFrame(resp_json, schema=_schema).sink_parquet(output_pq_path)
return output_pq_path
except Exception as exc:
print(f"Failed to download airline data {str(exc)} ")
raise
The above error handling is simply lazy and has no place in production. Here are a few problems.
- Type checking helps us guarantee that
output_diris actually aPath, in which case thejoinpathcannot fail and should not be in atryblock. The consequence is that our error message cannot include the file name/path we are attempting to write becauseoutput_pq_pathmight be unbound. - We cannot distinguish easily between IO errors creating the directory and writing the file without looking for string patterns in the error message. This is not very reliable, as error messages are not typically part of the public API/contract of packages. Exception types typically are more stable.
Bloated try-except blocks hinder useful error messages because such blocks
typically leave the except blocks with too many possibly unbounded local
variables.
Example
# /// script
# requires-python = ">=3.11"
# dependencies = ["polars", "requests"]
# ///
from pathlib import Path
from datetime import date
import requests
import polars as pl
def etl_airline_data_to_parquet(
base_url: str,
output_dir: Path,
publish_date: date,
) -> Path | None:
"""Download data from airline API to hive-style partitioned parquet for
a date that has not been processed yet, returning the newly written parquet
file if it was written, otherwise None if data was already processed/written"""
output_pq_path = output_dir.joinpath(
f"publish_date={publish_date.isoformat()}", "data.parquet"
)
if output_pq_path.is_file():
return None
try:
resp = requests.get(
base_url,
params={"publish_date": publish_date.isoformat()},
timeout=(5.0, 5.0),
)
except requests.ConnectionError as err:
_msg = f"Failed to connect to {base_url}. Check DNS resolution"
raise RuntimeError(_msg) from err
resp.raise_for_status()
resp_json = resp.json()
try:
output_pq_path.parent.mkdir(exist_ok=True, parents=True)
except OSError as err:
_msg = f"""Failed to create directory for {output_pq_path.parent}.
Check directory permissions"""
raise OSError(_msg) from err
_schema = {"flight_id": pl.Int64, "passengers": pl.Int64, "revenue": pl.Int64}
try:
dt_lzdf = pl.LazyFrame(resp_json, schema=_schema)
except (ValueError, TypeError) as err:
_msg = f"""Failed to coerce data to schema when processing airline for
url={base_url} and publish_date={publish_date.isoformat()}. Did the
alternative data team change their API again?"""
raise ValueError(_msg) from err
try:
dt_lzdf.sink_parquet(output_pq_path)
except OSError as err:
_msg = f"""Failed to write parquet file {output_pq_path} . The
filesystem might be full"""
raise OSError(_msg) from err
return output_pq_path
The above example does a better job of making each except block target a very
specific case, rather than just “something went wrong.” In certain long-running
services, this kind of structure opens the door to giving much better error
messages and passing that information up the call stack to recover, possibly by
returning an error value/enum variant everywhere this function raises.
Short of an errors-as-values approach, the purpose of re-raiseing is to include
information that would be lost in the call stack. For example, the error message on
requests.get would not include the query parameters, and so even if we log
the exception stack trace, we would never see exactly which request gave the
exception. Other error messages contain contextual information fundamentally
not found in the code and that can be invaluable for whomever is called in to
respond to this error state at 2:00AM.
Exceptions: On the other hand, any error state where the only correct action
is to crash the program and where the exception message already contains the
necessary information for debugging doesn’t need an explicit try-except.
Err.3. Consider errors-as-values
Iteration
It.1. Separate reused data generation from data collection with generator functions
Reason: Separation of “data generating logic” and collection into a data
structure enables easier testing paths of both components and enables easier
swapping of the target data structure, for example, from a list/tuple into a
set/frozenset. Moreover collecting an iterator is frequently both faster and
more memory efficient than repeated appends/inserts into a data structure for
many reasons, including the following.
- Different “consumers” of the data generating process can short-circuit/stop early, or choose to process the data iteratively or in batches.
- Iterator collection pushes the actual iteration (repeated calls to
__next__) to C rather than Python - Even invoking the
.appendor.addis a dict lookup in Python, while the insertion path is optimized in the C layer of CPython. - Iterator collection involves better size hints that reduce memory allocations
when compared to calling
.insertor.addmany times.
Only break out an iterative process to the generator when doing so improves readability (if a nested loop warrants a comment, maybe a reified 3-4-word symbol/function name would improve readability), testability, or laziness. Many iterate+collect/accumulate/accrete to data structure operations – especially over small/well-controlled amounts of data – do not benefit from generator functions.
Example
from collections.abc import Iterator
from datetime import date, timedelta
def _saturdays_between(start: date, end: date) -> Iterator[date]:
"""Yield all saturdays between `start` and `end`, inclusive of both
endpoints"""
# .isoweekday() returns 6 for Saturday and 7 for Sunday
days_until_saturday = (6 - start.isoweekday()) % 7
first_saturday = start + timedelta(days=days_until_saturday)
cursor = first_saturday
while start <= cursor <= end:
yield cursor
cursor = cursor + timedelta(days=7)
def _saturdays_with_zero_traffic(
*,
start: date,
end: date,
n_requests_by_date: dict[date, int],
) -> tuple[date, ...]:
"""Returns dates that are saturday with no requests, assuming
`n_requests_by_date` only contains entries for days with requests"""
return tuple(
d for d in _saturdays_between(start, end) if n_requests_by_date.get(d, 0) == 0
)
Bad Example
from datetime import date, timedelta
def _saturdays_with_zero_traffic(
*,
start: date,
end: date,
n_requests_by_date: dict[date, int],
) -> tuple[date, ...]:
res: list[date] = []
days_until_saturday = (6 - start.isoweekday()) % 7
first_saturday = start + timedelta(days=days_until_saturday)
cursor = first_saturday
while start <= cursor <= end:
if n_requests_by_date.get(cursor, 0) == 0:
res.append(cursor)
cursor = cursor + timedelta(days=7)
return tuple(res)
A simple inline comprehension expression also serves this purpose.
Exception: Certain algorithms may fit more naturally in one functional or logical scope.
Type Annotations
Type annotations and effective names for symbols (functions, classes, and variable names) remove the need for the vast majority of prose and structured/numpydoc-style documentation. When viewed as “statically analyzable documentation,” type annotations are well worth the verbosity, especially for programs whose reliability and maintainability. Changing a 1000+ line code base without any type annotations is a difficult endeavor, riddled with subtle traps. Especially past 5000 source lines of code, the local reasoning enabled by type annotations greatly aids maintainability. Programs that are strictly/strongly type annotated have a better path toward migration to a faster language, such as Golang because such programs facilitate local reasoning and make most details of the program explicit.
Effective type-checking involves the following.
- Sparing usage of
typing.Any, mainly near some external boundary - Typed function definitions in owned production code, especially function parameters and return types
- Invoke type checking in CI
Type.1. All production grade Python programs that value reproducibility, reliability, and maintainability should use the strictest possible type annotations and incorporate a type checker – such as mypy – in its strictest mode in CI and developer workflows.
Reason: Type checkers catch a wide variety of bugs before they show up in production. Type annotations enable contributors to understand a section of code/function using “local reasoning” rather than needing to trace through 50 function calls to determine “what is in this object/dictionary.”
Example
def worst(*, val1, val2, val3): ...
def still_bad(*, val1: int, val2: dict, val3: str) -> str: ...
def good(*, val1: int, val2: dict[str, str], val3: str) -> str: ...
Type.2. All services/applications/libraries that use boto3 for Amazon Web Services (AWS) interactions should use the types-boto3 library or its async sister types-aioboto3 for type checking and analysis.
The ubiquity of AWS in many tech stacks and the complexity of the SDK warrant
specific call outs. The AWS Python SDK parses schema definition files shared by
all AWS language-specific SDKs and generates classes and functions at runtime.
This invalidates most kinds of type inference that type-checkers such as mypy
can perform. One of the few answers to such an approach is code-generating the
type annotations from the same schema definition files, which is exactly the
approach of these two packages.
Keep the stub version aligned with the deployed aioboto3 version. Prefer
narrowly scoped extras such as types-aioboto3[s3,sqs] unless broad coverage is
justified. Treat type stubs as development dependencies unless they are
explicitly imported at runtime.
Example
# /// script
# requires-python = ">=3.11"
# dependencies = ["aioboto3", "types-aiobotocore[s3]", "types-aioboto3"]
# ///
from __future__ import annotations
from enum import unique, Enum
from typing import TYPE_CHECKING
import aioboto3
if TYPE_CHECKING:
from types_aiobotocore_s3.client import S3Client
@unique
class _GetOldestVersionIDErrorReason(Enum): ...
async def _s3_get_oldest_version_id(
bucket: str, key: str, s3_client: S3Client
) -> str | _GetOldestVersionIDErrorReason: ...
async def _near_script_entrypoint() -> None:
sess = aioboto3.Session()
async with sess.client("s3") as s3_client:
version_id_or_err = await _s3_get_oldest_version_id(
bucket="fake-bkt", key="fake-key", s3_client=s3_client
)
# Can reuse the client and session across async boundaries
...
Testing
TE.1. Use Pytest and write function-based pytest-style tests – avoid class-based unittest-style tests
Reason: Functions are flatter, simpler, and easier to compose. Pytest has a
better story for test setup and teardown. In Python, we do not use classes as
namespaces – we have modules for this purpose. unittest is useful within the
standard library for testing Python implementations themselves without needing
to make design decisions beyond “copy JUnit.” The decisions made by JUnit and
the Python standard library’s unittest do not reflect the constraints of most properly packaged Python programs. pytest, is designed to allow for concise,
to-the-point, maintainable tests that use the best of what Python has to offer,
especially test context/fixtures via context managers. Moreover the pytest CLI
offers extensive functionality, from surgical test selection, verbosity
configuration, reporting configuration, and debugger integration. The pytest
extension ecosystem is vast and formidable.
Example
from collections.abc import Iterator
from dataclasses import dataclass
import pytest
@dataclass(frozen=True)
class _TestDBContext: ...
@pytest.fixture
def seeded_empty_unauthenticated_database() -> Iterator[_TestDBContext]:
with _seeded_empty_unauthenticated_database_impl() as test_db_context:
yield test_db_context
def test_my_application_behavior(
seeded_empty_unauthenticated_database: _TestDBContext,
) -> None: ...
def test_another_application_behavior(
seeded_empty_unauthenticated_database: _TestDBContext,
) -> None: ...
Bad Example
class TestGrouping:
def setUp(self) -> None: ...
def tearDown(self) -> None: ...
def test_my_application_behavior(self) -> None: ...
def test_another_application_behavior(self) -> None: ...
Exceptions Only write class-based unittest-style tests if the existing
code base is using unittest-style tests.
TE.2. Prefer dependency injection and ‘interface-compatible’ implementations over mocking.
“Dependency injection” most commonly involves giving functions more parameters.
Reason: Mocking stands in the way of fearless refactoring and changing a program – for both people and LLMs. Every mock represents untested application behavior. That can be fine for situations/scenarios that are difficult to set up/hermetically reproduce, especially those related to an external service or flakiness, but dependency injection at least tests the “happy path” of a program more thoroughly than mocking does.
Passing constants as function parameters almost always invalidates the need to
mock them. Prefer to hoist the base URL of the HTTP service over mocking http
clients. The tmp_path pytest fixture should remove the need to mock any
filesystem operations.
Bad Example
# /// script
# requires-python = ">=3.11"
# dependencies = ["aiohttp", "pytest", "pytest-asyncio"]
# ///
"""Download information of top news stories from an API to the filesystem as
structured parquet"""
from datetime import date
from typing import NamedTuple
from pathlib import Path
import aiohttp
# app/library/cli/service module
class _NewsApiResponse(NamedTuple):
article_ids: tuple[int, ...]
_OUTPUT_PARQUET_DIR = "/mnt/big_data_prod/data/reputable_news/top_stories/parquet"
async def _fetch_top_news_events_of_day(
publish_date: date, top_k: int = 10
) -> _NewsApiResponse:
params = {"publish_date": publish_date.isoformat(), "top_k": top_k}
async with aiohttp.ClientSession() as session:
async with session.get(
"/api.reputable_news.xxyyzz/articles/top",
params=params,
) as response:
return _NewsApiResponse(tuple((await response.json())["article_ids"]))
async def etl_top_k_news_to_parquet(
publish_date: date,
) -> None:
records = await _fetch_top_news_events_of_day(publish_date)
# write parquet to `_OUTPUT_PARQUET_DIR` with pyarrow
...
# Test module
from unittest.mock import patch
import pytest
@pytest.mark.asyncio
async def test_etl_top_k_news_to_parquet(tmp_path: Path) -> None:
output_parquet_base_dir = tmp_path.joinpath("outpq")
output_parquet_base_dir.mkdir()
with (
patch(
"some.package._OUTPUT_PARQUET_DIR",
str(output_parquet_base_dir.absolute()),
),
patch("some.package.aiohttp.ClientSession", ...),
):
await etl_top_k_news_to_parquet(date.today())
assert ... # data `some_other_directory_in_ci` matches expectations
There are many issues with the above test.
- If we switch HTTP clients, our mock also needs to change.
- The test doesn’t actually verify that we invoke the HTTP client properly – the mock hides any issues passing headers, data, auth, etc. into the HTTP client.
- The test needs to know more about the implementation details than merely the ‘public’ interface.
We can fix all of these by hoisting some function parameters.
Example
# /// script
# requires-python = ">=3.11"
# dependencies = ["aiohttp", "pytest", "pytest-httpserver", "pytest-asyncio"]
# ///
from dataclasses import dataclass
from datetime import date
from typing import Final, NamedTuple
from pathlib import Path
import aiohttp
# app/library/cli/service
class _NewsApiResponse(NamedTuple):
article_ids: tuple[int, ...]
_PROD_OUTPUT_PARQUET_DIR: Final[Path] = Path(
"/mnt/big_data_prod/data/reputable_news/top_stories/parquet"
)
@dataclass(frozen=True)
class _TopArticlesRequestCtx:
timeout: float
session: aiohttp.ClientSession
...
@dataclass(frozen=True)
class _TopArticlesRequestParams:
top_k: int
publish_date: date
...
def as_params(self) -> dict[str, str]:
return {
"top_k": str(self.top_k),
"publish_date": self.publish_date.isoformat(),
# ...
}
async def _fetch_top_news_events_of_day(
req_params: _TopArticlesRequestParams, ctx: _TopArticlesRequestCtx
) -> _NewsApiResponse:
async with ctx.session.get(
"/articles/top",
params=req_params.as_params(),
timeout=ctx.timeout,
) as response:
response.raise_for_status()
resp_json = await response.json()
return _NewsApiResponse(tuple(resp_json["article_ids"]))
async def etl_top_k_news_to_parquet(
req_params: _TopArticlesRequestParams,
ctx: _TopArticlesRequestCtx,
output_parquet_base_dir: Path = _PROD_OUTPUT_PARQUET_DIR,
) -> None:
records = await _fetch_top_news_events_of_day(req_params, ctx)
# write as parquet to `output_parquet_base_dir` using pyarrow
...
# Test module
from pytest_httpserver import HTTPServer
import pytest
from pytest_httpserver import RequestHandler
# from my_etl_script import _TopArticlesRequestCtx, _TopArticlesRequestParams
@pytest.mark.asyncio
async def test_etl_top_k_news_to_parquet(
tmp_path: Path, httpserver: HTTPServer
) -> None:
fixed_date = date.fromisoformat("2026-01-01")
response_data = {"article_ids": list(range(10))}
httpserver.expect_request("/articles/top", method="GET").respond_with_json(
response_data
)
output_parquet_base_dir = tmp_path.joinpath("outpq")
output_parquet_base_dir.mkdir()
news_base_url = httpserver.url_for("/")
async with aiohttp.ClientSession(base_url=news_base_url) as session:
ctx = _TopArticlesRequestCtx(...)
req_params = _TopArticlesRequestParams(...)
await etl_top_k_news_to_parquet(req_params, ctx, output_parquet_base_dir)
assert ... # parquet data in `output_parquet_base_dir` matches expectations
assert len(httpserver.log) == 1, "We can't flood the service with requests"
The key here is to design function/API boundaries that enable dependency injection, in this case for the API URL and the target output directory.
While spinning up and programming an HTTP server that “minimally” reproduces the target external service behavior may not enable testing all of the idiosyncracies of the service and arguably may not be too structurally different from mocking the HTTP Client, the “hitting an HTTP service that is compatible enough with the ‘real’ external service” strategy at least allows verification of the proper invocation of the HTTP client. Standing up an HTTP server just for tests also enables switching the HTTP client implementation without changing tests. In that sense, this strategy “mocks” a more fundamental/intrinsic boundary of the program: it mocks “the actual external service” rather than the HTTP client/response. Of course, we could also skip mocking entirely and just hit the live HTTP external service. This is better left to an integration test, however. If this HTTP service is internal to an enterprise, not all CI environments may be set up to route to the live service. Moreover, running mutating requests/RPC calls can have side effects on live services which may be undesirable, especially given that we expect that the program will have bugs until it passes CI with the rigorous test suite.
Exceptions: Certain scenarios encountered in production are difficult to reliably reproduce, especially in CI environments that do not permit containers. Examples include idiosyncratic SFTP configurations, or flakiness (e.g. ’the first 2 service calls fail due to server load or gateway issues, but the 3rd call succeeds with some backoff).
TE.3. Prefer to test edge-case behavior via the package’s public interface over private functions/implementation details.
“Public” interface can mean different things for different programs. For libraries, it means “the functions and invocations that callers will use.” For services (e.g. HTTP or GRPC), the exposed endpoints constitute the “public” interface. For an “offline” application such as program that drains a queue to a database, the guideline is less precise. While there is no “public interface,” there are core function boundaries as opposed to functions that are more implementation details.
Reason Tests of private implementation details stand in the way of fearless refactors. Good tests can guide the changing of implementation details.
Example
# /// script
# requires-python = ">=3.11"
# dependencies = ["fastapi", "httpx2", "pytest"]
# ///
# Pretend our app code is an another module
# from myapp import app
import pytest
from collections.abc import Iterator
import uuid
from fastapi.testclient import TestClient
@pytest.fixture
def app_client() -> Iterator[TestClient]:
with TestClient(app) as c:
yield c
@pytest.fixture
def admin_api_key() -> str: ...
@pytest.fixture
def username_password_of_registered_user_not_in_admin_group(
app_client: TestClient,
admin_api_key: str,
) -> tuple[str, str]:
username = uuid.uuid4().hex[0:8]
password = "fake_password"
resp = app_client.post(
"/api/v1/user",
data={"groups": [], "username": username, "password": password},
headers={"x-app-api-key": admin_api_key},
)
assert resp.status_code == 201
return username, password
@pytest.fixture
def registered_job_id(
app_client: TestClient,
admin_api_key: str,
) -> int:
username = uuid.uuid4().hex[0:8]
password = "fake_password"
resp = app_client.post(
"/api/v1/job",
json={"spec": {...}},
headers={"x-app-api-key": admin_api_key},
)
assert resp.status_code == 201
job_id = int(resp.json()["job_id"])
return job_id
def test_user_not_in_admin_group_cant_stop_job(
app_client: TestClient,
username_password_of_registered_user_not_in_admin_group: tuple[str, str],
registered_job_id: int,
) -> None:
username, password = username_password_of_registered_user_not_in_admin_group
resp = app_client.post(
"/api/v1/login", data={"username": username, "password": password}
)
assert resp.status_code == 200
cookies = resp.cookies
resp = app_client.delete(f"/api/v1/job/{registered_job_id}", cookies=cookies)
assert resp.status_code == 403
Exceptions: Functions whose execution context is difficult to reproduce in a testing/CI environment may be better served by targeted tests. Though they should have clearly outlined interface boundaries.
TE.4. Avoid conditional branching in test cases by splitting up each branch into different test cases
Reason: if-statements within test bodies can impede a reviewer’s efforts
to determine exactly what a given test verifies. There are many situations
where a conditionally-branched test may never take one of the branches. One can
track this by computing code coverage specifically for tests. But avoiding
branches guarantees that “every assert in the body of a properly constructed
test function is executed and guarantees something about a program.
Fundamentally, branchless test bodies increase the signal-to-noise ratio of any
given test.
Bad Example
# /// script
# requires-python = ">=3.11"
# dependencies = ["pytest"]
# ///
import re
from typing import Final
# Application/service/library module
FILE_NAME_PATTERN: Final[re.Pattern[str]] = re.compile(
r"(?P<iso3_country_code>[A-Z]{3})\-(?P<publish_year>\d{4})\.csv"
)
# Testing module
import pytest
@pytest.mark.parametrize(
("instr", "expected_captures"),
[
pytest.param(
"CHN-2025.csv",
{"iso3_country_code": "CHN", "publish_year": "2025"},
id="CHN-2025.csv",
),
pytest.param(
"IND-1980.csv",
{"iso3_country_code": "IND", "publish_year": "1980"},
id="IND-1980.csv",
),
pytest.param(
"INDIA-1980.csv",
None,
id="bad_country_code",
),
pytest.param(
"FRA-123.csv",
None,
id="bad_publish_year",
),
],
)
def test_file_name_pattern_regex(
instr: str, expected_captures: dict[str, str] | None
) -> None:
got_mtch = FILE_NAME_PATTERN.fullmatch(instr)
if got_mtch:
assert expected_captures is not None
assert got_mtch.groupdict() == expected_captures
else:
assert expected_captures is None
In the above example, the urge to add some verification of the regex is reasonable, as regexes can be surprisingly complicated. However, the test body is more complex than it should be. We can fix this in one of two ways. If we are not willing to change the app/service/library source code/function signature, we can split up the tests into “good cases” and “bad cases.”
Example:
# /// script
# requires-python = ">=3.11"
# dependencies = ["pytest"]
# ///
import pytest
@pytest.mark.parametrize(
("instr", "expected_captures"),
[
pytest.param(
"CHN-2025.csv",
{"iso3_country_code": "CHN", "publish_year": "2025"},
id="CHN-2025.csv",
),
pytest.param(
"IND-1980.csv",
{"iso3_country_code": "IND", "publish_year": "1980"},
id="IND-1980.csv",
),
],
)
def test_file_name_pattern_regex(
instr: str,
expected_captures: dict[str, str],
) -> None:
got_mtch = FILE_NAME_PATTERN.fullmatch(instr)
assert got_mtch is not None
assert got_mtch.groupdict() == expected_captures
@pytest.mark.parametrize(
("instr",),
[
pytest.param("FRA-123.csv", id="bad_publish_year"),
pytest.param("INDIA-1980.csv", id="bad_country_code"),
],
)
def test_file_name_pattern_does_not_match(instr: str) -> None:
assert FILE_NAME_PATTERN.fullmatch(instr) is None
One can also change the application/service/library code to return more information that collapses multiple conditional branches as follows.
Example
from dataclasses import dataclass
import re
from typing import Final
# Application/service/library module
@dataclass(frozen=True)
class FilenameComponents:
iso3_country_code: str
publish_year: int
_FILE_NAME_PATTERN: Final[re.Pattern[str]] = re.compile(
r"(?P<iso3_country_code>[A-Z]{3})\-(?P<publish_year>\d{4})\.csv"
)
def parse_file_name(filename: str) -> FilenameComponents | None:
mtch = _FILE_NAME_PATTERN.fullmatch(filename)
if mtch is None:
return None
captured = mtch.groupdict()
# These hard accesses and parses should be safe when the regex matches
return FilenameComponents(
iso3_country_code=captured["iso3_country_code"],
publish_year=int(captured["publish_year"]),
)
# Test module
import pytest
@pytest.mark.parametrize(
("instr", "expected_parsed"),
[
pytest.param(
"CHN-2025.csv",
FilenameComponents(iso3_country_code="CHN", publish_year=2025),
id="CHN-2025.csv",
),
pytest.param(
"IND-1980.csv",
FilenameComponents(iso3_country_code="IND", publish_year=1980),
id="IND-1980.csv",
),
pytest.param(
"INDIA-1980.csv",
None,
id="bad_country_code",
),
pytest.param(
"FRA-123.csv",
None,
id="bad_publish_year",
),
],
)
def test_file_name_pattern_regex(
instr: str, expected_parsed: FilenameComponents | None
) -> None:
got = parse_file_name(instr)
assert got == expected_parsed
Logging
Log.1. Log One Canonical, Structured, Wide Completion/Outcome Event Per Unit of Work
Accumulate context over a meaningful unit of work and emit a comprehensive structured event, possibly with isolated logs for important events, such as security-relevant events, important state transitions, heartbeats/progress updates for especially long-running operations, or circuit-breakers. This applies for long-running HTTP/gRPC services, scripts, background jobs, queue workers, and workflows.
This is especially useful in programs that will run for longer than 10 seconds or programs that will be deployed to production because we typically ask “what is x program doing now/what is taking so long” of those such programs.
Reason: In the context of a long-running backend service that provides real
value in a business setting and that requires debugging during outages,
structured canonical events fundamentally transform debugging from “trudging
through swamps of text” to “answering directed questions with queries over
structured data.” The 8-32 fields emitted in a wide structured log event,
especially high-cardinality non-sensitive fields (e.g. n_rows_processed,
total_postgres_users_query_time_ms, n_s3_objects_put, job_status,
bytes_written_to_s3, subscription_tier, job_id, route,
status_code_detailed, environment) at the end of each unit of work overlap
tremendously with the (typically lower cardinality) fields that are the
bread-and-butter of dedicated metrics systems such as Datadog, the Grafana LGTM
Stack (Loki, Grafana, Tempo, Mimir) , VictoriaLogs, and Splunk. Machine-parsable
logs feed nicely into metrics systems and provide a quick path to more dedicated
metrics-based observability systems (e.g. Datadog logs vs metrics). In
production environments, logs have costs. Log storage costs money. Indexing for
searchability costs money. Noisy logs cost time when debugging. Extra downtime
from trudging through log noise costs money. Each log line must pull its
weight and actively help in some reasonably probable production outage scenario.
Example
# /// script
# requires-python = ">=3.11"
# dependencies = ["fastapi"]
# ///
import asyncio
from dataclasses import dataclass
from http import HTTPStatus
import logging
from enum import unique, StrEnum
from typing import Literal
from fastapi import APIRouter, FastAPI, Request, Response
from pydantic import BaseModel
_LOGGER = logging.getLogger(__name__)
_RawWideEventT = dict[str, str | int] # passed into _LOGGER.info(..., extra=...)
@dataclass(frozen=True)
class _GetJobSpecLogInfo:
"""Log-worthy information accreted while processing a GET job spec request"""
...
@unique
class _GetJobSpecificationErrorReason(StrEnum):
JOB_NOT_FOUND = "JOB_NOT_FOUND"
TIMEOUT = "TIMEOUT"
...
@dataclass(frozen=True)
class _GetJobSpecificationWideEventSuccess:
request_duration_ms: int
status: Literal["success"] = "success"
...
def as_raw(self) -> _RawWideEventT: ...
@dataclass(frozen=True)
class _GetJobSpecificationWideEventError:
request_duration_ms: int
status_code_detailed: _GetJobSpecificationErrorReason
error_message: str
status: Literal["error"] = "error"
...
def as_raw(self) -> _RawWideEventT: ...
_GetJobSpecificationWideEvent = (
_GetJobSpecificationWideEventSuccess | _GetJobSpecificationWideEventError
)
class JobSpecificationResponse(BaseModel): ...
api_v1_router = APIRouter(prefix="/api/v1")
app = FastAPI()
async def _get_job_spec_by_id(
job_id: int,
) -> tuple[JobSpecificationResponse, _GetJobSpecLogInfo]:
"""Query the postgresql database for a job spec"""
def _basic_wide_event(
ctx: _GetJobSpecLogInfo, request: Request
) -> _GetJobSpecificationWideEventSuccess:
"""Canonical event metadata for a success getting a job spec. Just moving
data around, and should never raise an error"""
...
def _get_job_uncaught_exception_wide_event(
job_id: int, request: Request
) -> _GetJobSpecificationWideEventError:
"""Canonical wide event metadata for an unexpected error getting a job spec"""
...
def _get_job_timeout_wide_event(
job_id: int, request: Request
) -> _GetJobSpecificationWideEventError:
"""Canonical wide event metadata for a timeout error getting a job spec"""
...
def _log_and_publish_metric(event: _GetJobSpecificationWideEvent) -> None: ...
@api_v1_router.get(
"/job/{job_id}", response_model=JobSpecificationResponse
) # Add more params for documentation as needed
async def _get_job_id(job_id: int, request: Request):
"""Get a job specification by its ID"""
try:
resp, ctx = await asyncio.wait_for(_get_job_spec_by_id(job_id), timeout=5.0)
except asyncio.TimeoutError:
_event = _get_job_timeout_wide_event(job_id, request)
_log_and_publish_metric(_event)
return Response(status_code=HTTPStatus.REQUEST_TIMEOUT)
except Exception:
_event = _get_job_uncaught_exception_wide_event(job_id, request)
_log_and_publish_metric(_event)
return Response(status_code=HTTPStatus.INTERNAL_SERVER_ERROR)
_event = _basic_wide_event(ctx, request)
_log_and_publish_metric(_event)
return resp
app.include_router(api_v1_router)
Any service worthy of existence and whose absence will be missed should be able to give some indication of its state per some unit of work that can be exposed at least via logs.
Command Line Interfaces
CLI.1. Use the standard library’s argparse module for command line interfaces.
Reason: The argparse module conforms with most unix-style CLI conventions
so that CLIs “feel” at home in unix-like environments, has very effective error
handling, gets regular improvements – including more humane error
handling and
color output introduced
in Python 3.14 –, and is featureful enough for the vast majority of CLIs. Most
CLIs should be simple and argparse tastefully implements all of the core
components of simple CLIs.
Example
This example is slightly more than minimal precisely to demonstrate how to integrate a CLI into, for example, a data processing operation.
#!/usr/bin/env python3
"""An example showing common features of the _many_ Python CLIs that I
write."""
from __future__ import annotations
import logging
from collections.abc import Sequence
from argparse import ArgumentParser, Namespace, ArgumentDefaultsHelpFormatter
from pathlib import Path
import sys
from datetime import date, timedelta
LOGGER = logging.getLogger(__name__)
_VERBOSITY_TO_LOG_LEVEL: dict[int, int] = {
0: logging.WARNING,
1: logging.INFO,
2: logging.DEBUG,
}
class _MyProgOpts(Namespace):
verbosity: int
start_date: date
end_date: date
input_data_base_dir: Path
output_parquet_base_dir: Path
def _get_parser() -> ArgumentParser:
parser = ArgumentParser(
description=__doc__, formatter_class=ArgumentDefaultsHelpFormatter
)
parser.add_argument(
"-v",
"--verbose",
action="count",
default=0,
help="Logging verbosity. Provide 0 to 2 times.",
dest="verbosity",
)
parser.add_argument(
"--start-date",
type=date.fromisoformat,
default=date.today() - timedelta(days=3),
help="The first date (inclusive) of data to process. ISO-8601 date format",
)
parser.add_argument(
"--end-date",
type=date.fromisoformat,
default=date.today(),
help="The last date (inclusive) of data to process. ISO-8601 date format",
)
parser.add_argument(
"input_data_base_dir",
type=Path,
help="The directory of the raw x dataset. zip archives named like 'some_data_2026-01-01.zip'",
)
parser.add_argument(
"output_parquet_base_dir",
type=Path,
help="The base directory of the hive-style partitioned parquet dataset to write/insert into",
)
return parser
def _run(opts: _MyProgOpts) -> int:
logging.basicConfig(
level=_VERBOSITY_TO_LOG_LEVEL.get(opts.verbosity, logging.DEBUG)
)
LOGGER.info("Beginning sample program with opts %s", opts)
from mypackage._workhorse import etl_data
written_paths = etl_data(
start=opts.start_date,
end=opts.end_date,
input_data_base_dir=opts.input_data_base_dir,
output_parquet_base_dir=opts.output_parquet_base_dir,
)
LOGGER.info(
"Wrote %s files to %s", len(written_paths), opts.output_parquet_base_dir
)
return 0
def main(args: Sequence[str] | None = None) -> int:
parser = _get_parser()
opts: _MyProgOpts = parser.parse_args(args, _MyProgOpts())
return _run(opts)
if __name__ == "__main__":
sys.exit(main())
There is no need to comment the main, _run, _get_parser or the Namespace
implementing class here. These are idioms and well understood among
professionals who write reliable Python. The great deal of functionality in the
CLI justifies the length.
Each function has a very specific purpose. main is the core interface between
the (C)Python runtime and the OS, insofar as parse_args typically receives
None at runtime and pulls from sys.argv. main bridges the world of a
string list of arguments with the world of static analyzability. We pass args
as a parameter to main to enable easy testing of main – just have the test
import main and call main with the same CLI flags one would pass to the
program. The run function is the bridge between argparse and the core
behavior of our program. In larger multi-file applications, prefer to not make
the CLI entrypoint module much larger than this. There are likely more useful
interfaces/invocation methods such as a HTTP service, Apache Airflow DAG, GRPC
service, or even as a library.
Even though type checkers do not entirely validate the information-flow from
argparse.ArgumentParser to the subclassed instance of argparse.Namespace,
the proximity of those two within the same module typically enables simple
debugging.
Bad Example Using a third-party CLI package such as typer for a simple CLI
with fewer than 16 parameters/options is a bad idea. While typer may cut down
on a bit of code and titillate the senses of those who like the idea of exposing
a decorated function and driving behavior from type annotations, the supply
chain risks from extra dependencies that add no runtime or “user”-facing value
(many CLIs are predominantly used by computers rather than by humans) outweigh
the burden of maybe 5-10 more lines of code.
Exception: More involved terminal-based user interfaces (e.g. TUIs,
“terminal-widget”-style libraries) may benefit from a CLI that is not
argparse. “Flashy” UI is not the distinguishing factor of most CLIs, however.
Of course, many “single-use”/“throwaway” Python scripts do not need to abide by
this example. Maybe a good rule of thumb is “if you ever expect a human or LLM
to use/understand the CLI more than once, use argparse in the style exemplified
above.
Packaging Python Programs
Pkg.1. Manually setting the PYTHONPATH environment variable is ALWAYS suspect.
This entire section largely consists of better alternatives.
Reason: There is almost never a need to set PYTHONPATH for a properly
packaged Python program. Very few who write Python programs in professional
settings understand the module resolution rules of PYTHONPATH. Indeed there
are many “gotchas.”
Pkg.2. Use PEP-723-style inline metadata for single-file standalone scripts that only require a couple of dependencies
Reason: A dedicated task-runner such as pipx or uv run can reliably
execute a script with dependencies without necessarily needing a full package
structure. PEP-723 is a nice middle ground between a fully reproducible package
structure and non-reproducible programs that use third-party dependencies
without declaring them.
Pkg.3. Prefer to put packaging metadata and tool configuration in pyproject.toml over setup.cfg and setup.py.
Reason: pyproject.toml, initially specified in
PEP-518, refined in
PEP-621, and continually specified by the
Python Packaging
Authority,
works with many tools such as hatch, uv, and pdm. By contrast, setup.py and
setup.cfg are only a part of setuptools. There are situations where
setuptools’ tendency to put egg files and other metadata in the project and
source directories is undesirable. Moreover the ability to easily switch between
build backends by only changing the build-system section of pyproject.toml
setup.cfg format have formalized specifications.
Exceptions: Certain teams center deployment tooling around setup.py. In
such cases, keeping with team conventions allows the team to operate coherently
and uniformly, at least until those tools and standards have deliberately
considered the potential benefits of using pyproject.toml.
Pkg.4. Minimize runtime/deployed dependencies
Reason: Every dependency at runtime is a potential attack vector. 2026 saw many high-profile supply chain attacks, many of which operated on transitive dependencies and hijacked deployment pipelines without the knowledge of library maintainers. Dependencies that only operate in CI (e.g. linters), may be less likely to cause problems for deployed applications. The days of PyPI mainly consisting of good faith actors are long gone.
Examples
- Prefer the
zoneinfostandard library module to the olderpytzthird party package. Added in Python 3.9,zoneinfois now available in every currently supported Python version. Prefer to also depend on thetzdatapackage maintained under the Python umbrella/organization when deploying to minimal environments lacking the olson/IANA timezone database (e.g. distroless containers). - Prefer the standard library’s
datetimefunctionalities overdateutil. The standard library properly and reliably handles almost all use cases of this once venerable package. - Prefer the
dataclassesstandard library module over theattrsthird party package that inspired it for the vast majority of “plain-old-data” containers especially for data that is generated within the Python program, unless there is a very good reason. - Prefer the
jsonstandard library module over theorjsonthird party package unless the application has specific performance requirements and JSON (de)serialization is a measured bottleneck.
The previous examples, though they involve fairly stable and reliable third-party packages, exist to highlight the fact that smaller dependency graphs translate to faster and cheaper CI runs, smaller supply chain attack surfaces for deployed services, quicker ramp-up for newer contributors (Python developers are expected to understand the standard library), and a better maintenance story insofar as more eyeballs around the world will notice security issues in CPython than in some third party package.
Exceptions: At the same time, avoid contorting the standard library into
something outside of its intended scope. For example, the standard library’s
urllib.request is
sufficient for simple scripts that make one or two HTTP requests to known
‘well-behaved’ services, however long-running services that will make many
requests should consider a battle-tested third party HTTP client such as
aiohttp for a better path to handling
real-world considerations like connection pooling, redirects, authentication,
cookies, stream decompression, and multipart form uploads. The lack of a
production-grade HTTP client is a known limitation of the Python standard
library as of 2026. As another example, database clients typically implement
involved wire protocols (sometimes wrapping high-performance C, C++, or Rust
implementations) that no HTTP or gRPC service should reimplement.
Pkg.5. Pin dependencies for long-running services
While libraries must permit a range of dependencies to prevent conflicts of consumer applications, the ’leaves’ in a dependency graph can and should pin dependencies, while making dependency upgrades dedicated commits.
Reason: This helps minimize supply chain attacks. A minor bug fix must not introduce vulnerabilities far beyond the scope of the affected code. Pinning dependencies cuts off many attack vectors and makes the risk explicit with dedicated dependency bump commits.
This can be accomplished in a few ways. For example, one can have a
requirements-minimal.txt, and generate frozen/pinned requirements.txt with
uv pip compile/pip-compile that is the actual list of dependencies used by
the service. uv.lock can also accomplish this, while also enforcing package
checksums and keep track of package source registries.
Polymorphism in Python
Poly.1. Prefer functools.singledispatch over inheritance for “closed” polymorphism when the input and output types are well-specified, when the different strategies are determined by externally defined specifications, and when one module/file can control all strategies/implementations.
Reason: This cleanly separates “behavior” from “data” without all of the
extra baggage of inheritance, especially when disallowing subclassing of
strategies/variants. Placing all implementations/registered functions in one
module minimizes the complexity of function registry via import side effect.
singledispatch tends to outlive its usefulness when there are enough
implementations to render organizing them into 1 module impractical.
functools.singledispatch implements a multi-method by dispatching on the type
of the first function parameter.
Example
from dataclasses import dataclass
@dataclass(frozen=True)
class _PreprocessContext: ...
# spec.py, possibly in another package/part of the monorepo
@dataclass(frozen=True)
class PDFAttachmentPreprocess:
compression_level: int
@dataclass(frozen=True)
class JpegAttachmentPreprocess:
max_target_size_bytes: int
@dataclass(frozen=True)
class ZipArchivePreprocess: ...
TAttachment = bytes
Preprocess = PDFAttachmentPreprocess | JpegAttachmentPreprocess | ZipArchivePreprocess
def infer_preprocess_procedure(attachment: TAttachment) -> Preprocess | None:
"""Use libmagic to determine what kind of file we are dealing with"""
# Dedicated registry module with all implementations
from functools import singledispatch
@singledispatch
def preprocess_attachment_before_upload(
preprocess: Preprocess,
attachment: TAttachment,
ctx: _PreprocessContext,
) -> bytes:
raise NotImplementedError()
@preprocess_attachment_before_upload.register
def _(
preprocess: JpegAttachmentPreprocess,
attachment: TAttachment,
ctx: _PreprocessContext,
) -> bytes:
...
# Maybe each implementation keeps the core body in separate modules and this
# registry forwards parameters to the variant-specific modules, leaving this
# registry module as a unified entrypoint to multiple fully known and
# controlled backends
@preprocess_attachment_before_upload.register
def _(
preprocess: PDFAttachmentPreprocess,
attachment: TAttachment,
ctx: _PreprocessContext,
) -> bytes: ...
Now we can either test just the preprocess_attachment_before_upload function
or we can give the registered function an actual name and call those in tests,
though we should prefer to just use preprocess_attachment_before_upload
because the rest of the library/service/application code will call that.
For single-file scripts where all of the strategies are specified and
implemented in one module, an if-elif-else chain encapsulated in a function is
simpler and faster.
Performance Note: singledispatch has some performance overhead when
compared to inheritance, measured in single-digit microseconds. This type of
dispatch typically happens once or a couple of times per work unit, rather than
millions of times in a tight loop. Profile before determining that this matters.
And if it does matter, maybe Python is not the best language for the
service/application.
Poly.2. Prefer a plugin system via packaged entrypoints for “open” polymorphism where the library does not control all implementations.
TODO.
Style
Style.1. Drive the core behavior of an application or service via functions rather than by classes/methods
Reason: Classes are for storing data or state, and for enforcing invariants of encapsulated/hidden attributes. Functions are fundamentally simpler and “flatter” than classes insofar as one rarely considers the distinctions between state, value, identity when using functions.
Bad Example
class APIDownloader:
def __init__(self, api_key: str, url: str) -> None: ...
async def download_to_file(self, output_file_name: str) -> None: ...
There are two dead giveaways that this class is unnecessary. First, a class with
2 public/invoked methods – one of them being __init__ – should almost always be a
function. Second, a class whose name is a noun derived from a verb should likely
be a function.
Example
async def download_from_api_to_file(
*,
url: str,
api_key: str,
output_file_name: str,
) -> None: ...