SUBSTRATE — Sagar Tailor, home

02Current location: Engine level, APIx

02engine

CONTRACT 02APIx

A quality-adjusted airfare price index for India, and the auditable pipeline that produces it.

01Identification

Identification

My role

Data acquisition, ingestion, and the analysis tooling that verifies it

Domain
Official statistics
Level
02 · Engine
Stack
Python · pytest
Team
3 contributors
Built
September 2026
Sourced statements
17

02Context

Context

The problem

An airfare index is only as credible as its inputs, and airfare inputs are hostile: quoted prices move constantly, carriers restrict automated access, and a single undocumented collection decision silently invalidates a longitudinal series.

Why it matters

Any index can produce a number. The question a statistical office would ask is whether you can reconstruct, months later, exactly which observations went in and which were excluded, and why.

03What it does

What it does

Built against MoSPI problem statement 26056. The engine is a frozen-specification airfare index; the harder half is the evidence pipeline underneath it, which has to collect real market observations under a fixed protocol and be able to prove afterwards what it collected and what it excluded.

I own the boundary where the system meets untrusted external data: the collectors, the store, the scheduler, the manual and automated collection paths, and the analysis tooling that replays exclusions and checks the result. Roughly 18,900 lines of Python across 72 files, of which about 5,100 lines are tests.

And where that is proven

Sole author of the ingestion layer — the observation store, the collector runner, the IndiGo parser and the scheduler — which is the boundary where the system meets untrusted external market data.
src/apix/ingestion/store.pyVerify: open src/apix/ingestion/store.py on GitHub in a new tab
The collection contract is frozen in the repository before data is gathered, so the protocol cannot be adjusted after seeing the results.
src/apix/ingestion/collectors/runner.pyVerify: open src/apix/ingestion/collectors/runner.py on GitHub in a new tab
Roughly 5,100 lines of the project's tests are mine, written against an external data source that cannot be relied upon to behave.
tests/test_collector_runner.pyVerify: open tests/test_collector_runner.py on GitHub in a new tab
The project publishes no index value, and says so in its own README, because the longitudinal evidence its frozen methodology requires does not exist yet.
README.mdVerify: open README.md on GitHub in a new tab

04Architecture

Architecture

A collection boundary, an evidence store, and a deterministic statistical spine that is forbidden from importing anything above it. My half is everything up to and including the store — the part that decides what counts as admissible evidence and can prove, months later, what it collected and what it refused.

Layers, shallow to deep

  1. Collection contractFrozen advance-purchase bands, lead-time windows and eligibility rules, committed before any data is gathered so the protocol cannot be adjusted after seeing results.
  2. Compliance gateClearance for a live run is required before a single request leaves the machine, and refused by default.
  3. Collectors and schedulerOne paced search per travel date. A site error gets one retry; an access challenge stops that source for the rest of the run.
  4. CanonicalisationEach search resolves to an Observation, an UnpricedFlight, or an exclusion carrying its reason. Nothing is discarded silently.
  5. Evidence storeFour SQLite tables — run, attempt, canonical observation, unpriced flight — beside a content-addressed directory of the raw bytes each one came from.
  6. Statistical spineJevons at the elementary level, Young / Modified Laspeyres above it, and the publication guards. Not my work.

What happens to one request

  1. Run config

    The frozen collection contract for this wave.

  2. Compliance gate

    Refuses before any request is made.

  3. Search

    Paced, per travel date, recorded whether or not it succeeds.

  4. Eligibility

    Earliest eligible flight per band, then the fare decision.

  5. Canonical record

    Observation, unpriced flight, or exclusion with a reason.

  6. Export and manifest

    Written first, so the evidence survives a refused load.

  7. Store and verify

    Loaded into the store, then read back and checked.

src/apix/ingestion/collectors/runner.pyVerify: open src/apix/ingestion/collectors/runner.py on GitHub in a new tab

05Decisions

Decisions

Each one states what forced it, what was rejected, the reasoning, and what it cost.

  1. 01Decision

    A single SQLite file beside a content-addressed directory of raw artifacts, using only the standard library.

    over A server database or a cloud object store.

    The pressure
    The specification requires that a publication can be re-run from its recorded version vector and compared bit for bit.
    Why this way
    Reproducibility is a property of being able to find the original bytes again, not of the database engine. A file can be copied, diffed, and have its checksum committed.
    What it cost and bought
    No operational cost and no new dependency, and PostgreSQL can replace the file later without touching a call site — the calls are the contract, not the storage.
  2. 02Decision

    Store no expected-cells table at all, and make publication take the denominator as a required argument with no default.

    over A schema field recording what the collector expected to find.

    The pressure
    Coverage has to be reported against the cells that were expected, and “expected cells” is undefined in the frozen specification.
    Why this way
    Storing the cells we saw and reading them back as the cells we expected is circular. Coverage measured against our own success can never fall, and would report 100% on a day the collector was blocked.
    What it cost and bought
    The ambiguity stays open and named rather than being quietly answered by a schema, and no number can be published until someone declares the denominator out loud.
  3. 03Decision

    Record an attempt for every search — including ones that failed, and ones never made because an earlier search hit an access challenge.

    over Persisting only the observations that were collected.

    The pressure
    Missing data is what an index is most easily wrong about, and a store of successful observations cannot describe it.
    Why this way
    Missingness is measurable only from a record of attempts. Without it, a blocked run and a quiet market produce the same rows.
    What it cost and bought
    The exclusion replay can be reconstructed long afterwards, and the two cases stay distinguishable.
  4. 04Decision

    Tag automated, sandbox and fixture runs in the frame identifier, and admit only primary frames to the index.

    over Letting automated runs feed the index once the adapter worked.

    The pressure
    An automated collector is the convenient path and also the easiest way to contaminate a frozen longitudinal protocol.
    Why this way
    Adopting automated collection as a production frame is an owner decision recorded in a decision record, not a default that arrives with a working adapter.
    What it cost and bought
    The browser path can be exercised freely without any run of it ever becoming index input.
  5. 05Decision

    Write the run export and manifest before handing anything to the store.

    over Exporting after a successful load.

    The pressure
    Evidence is written at the end of a run, which is exactly when a load can be refused.
    Why this way
    A refused load must not also destroy the record of what was collected — that is the run you most need the audit trail for.
    What it cost and bought
    The trail survives the failure, and a rejected wave can be examined instead of repeated blind.

06Challenges

Challenges

07Results

Results

Commits on main
16/29
55% of the project
Pull requests
18/32
17 merged
Lines of Python
18,900
Across 72 files
Lines of tests
5,100
Mine

08Attribution

Attribution

Three contributors. I am the largest by commits and own the acquisition boundary; the statistical core and the interface work are not mine.

Ingestion, collectors, scheduler, collection and analysis toolingSagar16 of 29 commits
Statistical core, methodology and index specificationSoumyadeb Tripathy8 of 29 commits
Interface, art direction and jury dashboard surfaceBasant Bhushan5 of 29 commits

Not mine:The index methodology and the statistical engine are Soumyadeb's work, and the presentation layer is largely Basant's. I did not design the index.

09Next iteration

Next iteration

Each of these is something the repository already records as unfinished, not a feature list. Where a project states no such thing, this section is absent rather than invented.

  1. 01Open

    Run the second collection wave, so the longitudinal step has a matched t and t−7 pair to compare.

    The engine is implemented and tested and still publishes nothing, because one wave produces zero matched pairs and the chaining guard refuses an index value without them. The next run is what unlocks the first real number.

  2. 02Open

    Establish a second source and a second route, so the panel stops being one carrier on one sector.

    National representativeness is recorded as not established. Deduplication and source precedence are implemented but degenerate on a single-source panel — exercised is not validated.

Still to write

Still to write

9 of 10 sections written and sourced

The rest are absent rather than filled with plausible prose, which is the whole point: inventing them would cost exactly the credibility the rest of this page is built to earn.

  • LessonsNot written yet
What these mean
Key to the states above.
Not written yet
Real and known, but not written up.

Navigation

Use arrow keys to select, Enter to open, Escape to close.