A financial data API serves 27,273 JSON files—equities, mutual funds, ETFs—to a compliance reporting system. The upstream data provider sends a fresh export every few days. Every export has problems: bare NaN values that crash JSON parsers, European comma decimals, missing fields, phantom peer references, duplicate records. Each bug causes an HTTP 500 in production.
The team needed a QA gate: a script that runs on the raw data before it enters the system, catches every category of defect, and reports what the upstream provider needs to fix.
This is the story of building that gate in a single day using Honest Code principles, and how each principle paid for itself in a concrete, measurable way.
The first instinct is to reach for a class. A RecordValidator with an __init__ that loads the schema, a validate() method that sets instance state, maybe a Report class to accumulate results.
Instead: a plain dict. Two hundred eighty-one key-value pairs, field name to type string.
EQUITY_SCHEMA: dict[str, str] = {
"id": "str", "symbol": "str", "name": "str",
"sector": "str", "industry": "str", "country": "str",
"market_cap": "num", "pe": "num", "ev_ebitda": "num",
"buckler_score": "num", "growth_score": "num",
# ... 281 fields total
}
No class hierarchy. No abstract base validator. No FieldType enum. Just a dict that says "these fields exist, and they're either strings or numbers." The check functions iterate over it. Adding a new field means adding one line.
Each check is a standalone function. Data goes in, a list of issues comes out. No self, no instance state, no side effects.
def check_score_range(data: dict, filename: str) -> list[dict]:
"""Scores should be 0-100."""
issues = []
for key in SCORE_FIELDS:
val = data.get(key)
if val is None or not isinstance(val, (int, float)):
continue
if val < 0 or val > 100:
issues.append({
"file": filename,
"check": "score_out_of_range",
"severity": "WARN",
"detail": f"'{key}' = {val} — expected 0-100.",
})
return issues
Every check follows the same shape: take a dict, return a list. No check knows about any other check. No check reads files. No check prints anything. The function is its own complete specification.
We wrote eleven of these in under an hour:
check_valid_json() # Bare NaN, comma decimals, parse errors check_filename() # UUID format validation check_id_matches_filename() # Internal ID must match the filename check_missing_keys() # All schema fields must be present check_types() # String fields hold strings, numbers hold numbers check_sentinel_values() # "--", "N/A", "null" should be actual null check_name_present() # Every security needs a name check_score_range() # Scores must be 0-100 check_zero_score_with_peers()# Zero score with nonzero peers = calculation failure check_duplicate_peers() # No repeated peers, no self-as-peer check_numeric_strings() # "14.89" should be 14.89
Testing any check is one line: assert check_score_range({"buckler_score": 150}, "test.json") == [expected_issue]. No mocks. No setup. No teardown.
How do you wire eleven independent checks together? Not with a base class and abstract methods. Not with a plugin registry. With one function that calls them in sequence:
def validate_record(raw: str, filename: str) -> list[dict]:
issues = []
issues.extend(check_filename(filename))
data, parse_issues = check_valid_json(raw, filename)
issues.extend(parse_issues)
if data is None:
return issues
issues.extend(check_id_matches_filename(data, filename))
issues.extend(check_name_present(data, filename))
issues.extend(check_sentinel_values(data, filename))
issues.extend(check_duplicate_peers(data, filename))
if is_equity(data):
issues.extend(check_missing_keys(data, filename, EQUITY_SCHEMA))
issues.extend(check_types(data, filename, EQUITY_SCHEMA))
issues.extend(check_score_range(data, filename))
return issues
Adding a new check is one line. Removing a check is deleting one line. The composition is visible at a glance—no method resolution order, no super() chains, no plugin discovery. A junior engineer can read this function and understand every check that runs, in what order, with what data.
Here's what main() looks like. Read files. Call pure functions. Write results.
def main():
# I/O: read all files from disk
records = read_all_files(data_dir)
# Pure: validate every record
per_record_issues = validate_all(records)
# Pure: build indexes for batch checks
indexes = build_indexes(records)
# I/O: load baseline from previous run
baseline = load_baseline()
# Pure: run batch checks
batch_issues = run_batch_checks(indexes, baseline)
# Pure: build structured report
report = build_report(all_issues, len(files), ...)
# I/O: write report, save baseline, print verdict
write_report(report)
save_baseline(indexes)
print_verdict(report)
Every line is either I/O or a pure function call. You can see the data flow. The pure functions in the middle don't know files exist. They don't know about baselines on disk. They don't know about stdout. They take data in and return data out.
This is not an aesthetic preference. It has an immediate, practical consequence.
I/O AT THE BOUNDARYThe script took 11.7 seconds to validate 27,273 files. We wanted it faster. Here's the question we asked:
validate_record() run in parallel?Because we followed Honest Code principles, the answer was immediately obvious. validate_record() is pure. It takes a string and a filename. It returns a list. It reads no files. It mutates no shared state. It has no side effects.
Of course it can run in parallel. It's embarrassingly parallel by construction.
If we had written a Validator class with instance state—a running issue count, a reference to a shared report object, a file handle for logging—parallelism would have required locks, queues, careful synchronization, and a rewrite. Instead:
# Before: sequential
results = [validate_record(raw, fname) for fname, raw in records]
# After: parallel across all CPU cores
with ProcessPoolExecutor() as pool:
results = list(pool.map(validate_chunk, chunks))
That's it. We split the records into chunks, farmed them out to worker processes, and collected the results. The pure functions didn't change at all. The only code that changed was in main()—at the I/O boundary.
We had to use ProcessPoolExecutor (separate processes) instead of ThreadPoolExecutor (threads) because Python's GIL prevents threads from running CPU-bound Python code in parallel. We benchmarked both:
Threads made it worse. Processes made it 2.6x faster. But neither change required touching the validation logic. The I/O boundary absorbed the change. The pure functions stayed pure.
PURE FUNCTIONS OVER METHODS I/O AT THE BOUNDARYBy the end of the day, the file had a clean three-layer structure:
┌─────────────────────────────────────────────────┐
│ SCHEMA DECLARATION │
│ EQUITY_SCHEMA = {"field": "type", ...} │
│ Just data. 281 fields. One dict. │
├─────────────────────────────────────────────────┤
│ PURE CHECK FUNCTIONS │
│ check_valid_json() → list[dict] │
│ check_score_range() → list[dict] │
│ check_dead_columns() → list[dict] │
│ ... 16 functions, all pure, all testable │
├─────────────────────────────────────────────────┤
│ COMPOSITION (DISPATCH TABLES) │
│ validate_record() — per-record dispatch │
│ run_batch_checks() — batch dispatch │
│ build_report() — pure report builder │
├─────────────────────────────────────────────────┤
│ I/O BOUNDARY (main only) │
│ Read files → pure functions → write results │
│ Parallelism lives here, not in the logic │
└─────────────────────────────────────────────────┘
On the first real run, the QA gate found:
884 files with bare NaN values (REJECT) 108 completely dead columns (REJECT) 1333 suspicious zero scores (WARN) 1225 peer name mismatches (WARN) 108 peer symbols not in dataset (WARN) 84 duplicate peer symbols (WARN) 36 duplicate ticker symbols (WARN) 6 sentinel placeholder strings (WARN)
The peer name mismatches revealed a systematic upstream bug: the data provider was matching peers by ticker symbol without considering the exchange. WFC on the Toronto Stock Exchange is Wall Financial Corporation. WFC on the New York Stock Exchange is Wells Fargo. The provider was crossing the streams.
This bug had been silently corrupting peer comparison data for months. A pure function found it in 1.35 seconds.
This is the entire process for adding a new validation rule:
def check_negative_market_cap(data: dict, filename: str) -> list[dict]:
val = data.get("market_cap")
if isinstance(val, (int, float)) and val < 0:
return [{"file": filename, "check": "negative_market_cap",
"severity": "REJECT",
"detail": f"market_cap is {val} — cannot be negative."}]
return []
def validate_record(raw, filename):
...
issues.extend(check_negative_market_cap(data, filename)) # ← this line
...
return issues
No base class to extend. No interface to implement. No configuration file to update. No registration mechanism. Write the function, add it to the list.
The schema is a plain dict. Issues are plain dicts. No __init__, no self, no inheritance. Adding a field or check means adding a line, not a class.
Every check is input-in, output-out. Testable with assert. No mocks needed. And because they're pure, they're automatically parallelizable.
File reads, printing, and report writing happen only in main(). This made parallelism a three-line change at the boundary instead of a rewrite of the logic.
validate_record() is a visible list of function calls. Adding a check is one line. No plugin system, no abstract methods, no hidden dispatch.
The schema dict declares truth. The functions enforce it. Changing what's valid means changing the declaration, not rewriting validation logic.
None of these principles felt like extra work while writing the code. Typed dicts are less code than classes. Pure functions are simpler than methods with state. Flat composition is more readable than inheritance hierarchies. I/O at the boundary is fewer lines than scattered file access.
The payoff came later—when we needed parallelism and it was three lines, when we needed a new check and it was one line, when we needed to understand the data flow and it was one function.
Honest Code isn't a tax on your productivity. It's an investment that compounds every time you touch the code again.