Auditing Collector-Generated Graduation Labels on Pump.fun: Measurement Error and Temporal Non-Generalization

By Arati Uday Kamat

Rating

1509
Battle Count: 68

Relevance

4/10
The paper is relevant to quantitative trading in the cryptocurrency/memecoin domain as a methodological cautionary study. It demonstrates that (1) collector-generated outcome labels on pump.fun cannot be equated with platform-side graduation events, (2) in-sample discrimination (AUROC 0.8594) can completely fail to generalize temporally (validation AUROC 0.4642, indistinguishable from chance), and (3) mcap functional form is unstable across specifications. These findings directly impact any quantitative strategy that relies on off-chain collector data for memecoin graduation prediction. However, the paper does not propose a trading strategy or predictive model that works; it documents a negative result. The methodological framework (pre-registration, stability gates, temporal holdout) is transferable to other crypto microstructure studies.

Implementation Complexity

6/10
The primary model is a standard binary logistic regression (GLM with Binomial/Logit link) implemented via statsmodels, which is straightforward. However, the full pipeline includes: (1) source-code audit of the V3 collector's polling mechanism, (2) reconstruction of collector state from polling logs, (3) pre-registered stability-gate framework with 9 automated evaluations across Gates B, C, D, G, (4) day-block bootstrap with cluster-robust covariance and careful handling of design_info reuse failures, (5) 17 pre-registered sensitivity analyses, (6) calibration diagnostics (slope, intercept, decile plots), (7) identity checks for AUROC computation, and (8) full provenance tracking with SHA-256 verification. The statistical methodology is moderate; the measurement audit and reproducibility infrastructure add significant complexity.

Reproducibility

5/5
Exceptionally reproducible: frozen inputs sealed under SHA-256 before modeling, pre-registration document with SHA-256 hash, full code and data archived at Zenodo (concept DOI 10.5281/zenodo.20633486), deterministic cohort flow ledger, locked package versions (numpy 2.5.1, pandas 2.3.3, scipy 1.18.0, statsmodels 0.14.6, patsy 1.0.2, scikit-learn 1.9.0), provenance manifest covering every deposited file, REBUILD_ANALYSIS.sh pipeline script, SHA256SUMS verification, and a fully documented correction history from v1.3 to v5. Failed bootstrap replicates are neither replaced nor rerolled. All sensitivity analyses are enumerated and their execution status explicitly reported.

About this paper

Methodology: Pre-registered binary logistic association model with temporal holdout validation and source-code measurement audit. Problem types: Classification, Measurement Error, Temporal Validation, Model Calibration.

The interactive Everscope explorer (charts, battles, favorites) loads below.