Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough

arXiv cs.LG 论文

摘要

This paper introduces a screen-and-confirm protocol to certify whether conditioning signals improve temporal point process models of customer return timing, finding that continuous-time decay is nearly sufficient and added conditioning is redundant or harmful.

arXiv:2608.11555v1 Announce Type: new Abstract: Practitioners enrich customer-return models with ever more signals (lifetime value, category, recency/frequency, calendar, geography), and the temporal-point-process (TPP) literature follows suit with covariate- and external-covariate-conditioned intensities. But does any of it improve the timing, and how would you know? A null ("feature X doesn't help") is only meaningful if the model could have found a signal. We make two contributions--a method and a measurement--to answer this credibly. (i) A screen-and-confirm protocol that certifies whether a candidate signal improves a TPP's event-timing likelihood: a positive control plants a coupling of known strength and confirms the model recovers it, so a real-data null can be read as "no signal" rather than "weak method." The control is validated for categorical and continuous encodings, and on a real clock-driven dataset (NYC taxi hour-of-day). (ii) A model-free ceiling quantifying how little of customer-return timing is point-predictable at all (a single-digit percentage of gap variance from any covariate; returns are near-memoryless). With these we certify a clean result on three public benchmarks (Amazon, Taobao, RetailRocket) and a real marketplace (Thumbtack): the inter-event clock--continuous-time decay, long known to beat frozen-intensity models--is nearly sufficient, and the conditioning the field keeps adding is redundant or harmful on top of it (statistically null on the public benchmarks, at most 0.06 NLL; null to mildly harmful on the marketplace). We do not claim to discover that decay helps; our contribution is the tools that turn "conditioning doesn't help" into a checkable, certified statement--plus an honest-evaluation account of the read-out/leakage pitfalls we hit and retracted.
查看原文
查看缓存全文

缓存时间: 2026/08/13 15:37

# Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough
Source: [https://arxiv.org/html/2608.11555](https://arxiv.org/html/2608.11555)
,Vineeth LoganathanAffiliation:Thumbtack, Inc\.,San Francisco,CA,USAemail:[vloganathan@thumbtack\.com](mailto:[email protected]),Shishir DashAffiliation:Thumbtack, Inc\.,San Francisco,CA,USAemail:[shishirdash@thumbtack\.com](mailto:[email protected])andVijay RaghavanAffiliation:Thumbtack, Inc\.,San Francisco,CA,USAemail:[vraghavan@thumbtack\.com](mailto:[email protected])

© none

###### Abstract\.

Practitioners enrich customer\-return models with ever more signals \(lifetime value, category, recency/frequency, calendar, geography\), and the temporal\-point\-process \(TPP\) literature follows suit with covariate\- and external\-covariate\-conditioned intensities\. But*does any of it improve the timing*, and how would you know? A null \(“featureXXdoesn’t help”\) is only meaningful if the model*could*have found a signal\. We make two contributions—a method and a measurement—to answer this credibly\.\(i\)Ascreen\-and\-confirm protocolthat*certifies*whether a candidate signal improves a TPP’s event\-timing likelihood: a positive control plants a coupling of known strength and confirms the model recovers it, so a real\-data null can be read as “no signal” rather than “weak method\.” The control is validated for categorical and continuous encodings, and on a real clock\-driven dataset \(NYC taxi hour\-of\-day\)\.\(ii\)Amodel\-free ceilingquantifying how little of customer\-return timing is*point*\-predictable at all \(≲5%\\lesssim\\\!5\\%of gap variance from any covariate; returns are near\-memoryless\)\. With these we certify a clean result on three public benchmarks \(Amazon, Taobao, RetailRocket\) and a real marketplace \(Thumbtack\): the inter\-event clock—continuous\-time*decay*, long known to beat frozen\-intensity models—is nearly*sufficient*, and the conditioning the field keeps adding is redundant or harmful on top of it \(statistically null on the public benchmarks,≲0\.06\\lesssim\\\!0\.06NLL; null to mildly harmful on the marketplace\)\. We do*not*claim to discover that decay helps; our contribution is the tools that turn “conditioning doesn’t help” into a checkable, certified statement—plus an honest\-evaluation account of the read\-out/leakage pitfalls we hit and retracted\.

###### Keywords:

temporal point processes, customer return, conditioning, positive controls, evaluation

††footnotetext:Accepted at the 5th Workshop on End\-to\-End Customer Journey Optimization \(KDD 2026\)\.## 1\.Introduction

“When will this customer come back?” underlies notification timing, CRM prioritization, and lifecycle modeling\. Temporal point processes \(TPPs\) answer it by modeling the conditional intensityλ∗​\(t∣ℋt\)\\lambda^\{\*\}\(t\\mid\\mathcal\{H\}\_\{t\}\)—the instantaneous event rate given the past\. Neural TPPs \(RMTPP\([rmtpp](https://arxiv.org/html/2608.11555#bib.bib1)\)→\\toNHP\([nhp](https://arxiv.org/html/2608.11555#bib.bib2)\)→\\toTransformer\-Hawkes/SAHP\([thp](https://arxiv.org/html/2608.11555#bib.bib3);[sahp](https://arxiv.org/html/2608.11555#bib.bib4)\)→\\tostate\-space models\([s2p2](https://arxiv.org/html/2608.11555#bib.bib9)\)\) have steadily improved likelihood fit\. But a practitioner choosing a model for customer return faces a question the literature does not answer cleanly:which components actually move return\-timing, and how would you know?Enriching event models with such signals is common practice: marked TPPs condition on event type/category by design\([rmtpp](https://arxiv.org/html/2608.11555#bib.bib1);[nhp](https://arxiv.org/html/2608.11555#bib.bib2)\); covariate TPPs add feature vectors and even learn their importance\([transfeat](https://arxiv.org/html/2608.11555#bib.bib13)\); external\-covariate TPPs inject seasonal/periodic drivers into the intensity\([metp](https://arxiv.org/html/2608.11555#bib.bib14)\); and recency/frequency \(RFM\) and value features power deep churn and lifetime\-value models\([ziln](https://arxiv.org/html/2608.11555#bib.bib17)\)\. These works typically*add*a signal and report a gain, in domains where it plausibly drives timing\. We ask a different question for customer\-*return*timing: not whether such a covariate*can*be added, but whether it adds anything once a strong inter\-event backbone \(decay\) is already in place—and how to tell a true null from a method too weak to find the signal\.

We answer empirically on three public benchmarks \(Amazon, Taobao, RetailRocket\) plus a real marketplace \(Thumbtack\), reporting temporal negative log\-likelihood \(NLL\), inter\-event RMSE, and MAE\. Our contributions:

- •C1\. A screen\-and\-confirm certification \(method\)\.A protocol that certifies whether a candidate signal improves a TPP’s event\-timing likelihood: a positive control plants a coupling of known strength and confirms recovery \(categorical*and*continuous encodings, plus a real clock\-driven dataset, NYC taxi\), so a real\-data null reads as “no signal,” not “weak method” \(§[5\.5](https://arxiv.org/html/2608.11555#S5.SS5)\)\.
- •C2\. A model\-free ceiling \(measurement\)\.How little of customer\-return timing is*point*\-predictable at all:≲5%\\lesssim\\\!5\\%of gap variance from any covariate; returns are near\-memoryless \(§[5\.4](https://arxiv.org/html/2608.11555#S5.SS4)\)\.
- •C3\. The certified result\.Using C1–C2, on three public benchmarks and a real marketplace: continuous\-time*decay*\(the inter\-event clock\) is nearly*sufficient*, and the conditioning practitioners keep adding—LTV, category, RFM, calendar, geography—is redundant or harmful on top of it \(statistically null on the public benchmarks,≲0\.06\\lesssim\\\!0\.06NLL; null\-to\-harmful, up to\+0\.65\+0\.65, on the marketplace\), for a*dual mechanism*reason \(§[5\.1](https://arxiv.org/html/2608.11555#S5.SS1), §[5\.3](https://arxiv.org/html/2608.11555#S5.SS3), §[6](https://arxiv.org/html/2608.11555#S6)\)\.

The two contributions are one story\.Our contribution is the*tools*—a screen\-and\-confirm certification \(C1\) and a model\-free ceiling \(C2\); decay is a known mechanism, not our finding\. The result \(C3, decay nearly suffices\) and the tools are inseparable: “conditioning doesn’t help” is trustworthy*only*because the positive control shows the null means “no signal,” not “a method too weak to find it\.” The tool is what licenses the result\. \(Plain THP, our frozen\-intensity reference, is an*ablation*, not a deployment target\.\)

The transferable lesson: for return\-*timing*, the inter\-event clock \(decay\) is the reliable lever, andbefore believing that a feature helps, screen it against a positive control—so that a null means “no signal in the data,” not “a method too weak to find it\.”

## 2\.Related Work

Neural TPP methods\.RMTPP\([rmtpp](https://arxiv.org/html/2608.11555#bib.bib1)\)→\\toNHP\([nhp](https://arxiv.org/html/2608.11555#bib.bib2)\)→\\toattention models THP\([thp](https://arxiv.org/html/2608.11555#bib.bib3)\)and SAHP\([sahp](https://arxiv.org/html/2608.11555#bib.bib4)\)→\\toAttNHP\([attnhp](https://arxiv.org/html/2608.11555#bib.bib5)\); intensity\-free density models\([intfree](https://arxiv.org/html/2608.11555#bib.bib6)\); recent continuous\-time state\-space / latent\-linear\-Hawkes models \(S2P2\([s2p2](https://arxiv.org/html/2608.11555#bib.bib9)\)\)\. EasyTPP\([easytpp](https://arxiv.org/html/2608.11555#bib.bib8)\)is the standard benchmark/codebase; the field evaluates on NLL, RMSE/MAE, and mark accuracy\([review](https://arxiv.org/html/2608.11555#bib.bib10)\)\. We compare these backbones plus conditioning components on the customer\-return task\.

Customer return / CLV\.Buy\-till\-you\-die models \(BG/NBD\([bgnbd](https://arxiv.org/html/2608.11555#bib.bib15)\); Pareto/NBD\([paretonbd](https://arxiv.org/html/2608.11555#bib.bib16)\)\) and deep CLV\([ziln](https://arxiv.org/html/2608.11555#bib.bib17)\)model value/return but not the full event\-time intensity\. Notably, BTYD models*assume*Poisson purchasing while alive—exponential, memoryless inter\-purchase gaps per customer; our model\-free ceiling \(§[5\.4](https://arxiv.org/html/2608.11555#S5.SS4)\) independently confirms that forty\-year\-old assumption on four datasets, and the small residual previous\-gap regularity \(r2r^\{2\}11–6%6\\%\) is the mild timing regularity the Pareto/GGG extension models\([paretoggg](https://arxiv.org/html/2608.11555#bib.bib22)\)\. Hazard/point\-process treatments of user return include Kapoor et al\.\([kapoor](https://arxiv.org/html/2608.11555#bib.bib19)\)and Du et al\.\([du2015](https://arxiv.org/html/2608.11555#bib.bib20)\); the closest neural treatment is the RNN survival model of Grob et al\.\([grob](https://arxiv.org/html/2608.11555#bib.bib11)\), which predicts a*single*next return time per user\. We instead model the full intensityλ∗\\lambda^\{\*\}and ask*which component matters on what data*, varying the backbone \(NHP, THP, S2P2\) and conditioning rather than committing to one architecture\.

Conditioning TPPs on covariates\.Injecting side information into the intensity is well studied: covariate TPPs encode feature vectors and learn per\-feature importance \(TransFeat\-TPP\([transfeat](https://arxiv.org/html/2608.11555#bib.bib13)\)\); external\-covariate TPPs decompose seasonal/periodic drivers across temporal granularities \(METP\([metp](https://arxiv.org/html/2608.11555#bib.bib14)\)\); and marked TPPs\([rmtpp](https://arxiv.org/html/2608.11555#bib.bib1);[thp](https://arxiv.org/html/2608.11555#bib.bib3)\)condition on event type by construction\. These report accuracy gains in settings where the covariate drives event dynamics\. We do not propose a new conditioning mechanism; we measure whether the standard ones help for customer\-return*timing*against a decay backbone, and add a positive\-control screen so a null is interpretable\.

Exogenous and causal TPPs\.Estimating whether an*external*signal \(calendar, geography, or an intervention such as marketing\) drives event timing is, in general, a causal question; counterfactual/causal TPPs\([counterfactualtpp](https://arxiv.org/html/2608.11555#bib.bib12)\)formalize it but require interventional or carefully\-controlled data\. Our screen\-and\-confirm test \(§[5\.5](https://arxiv.org/html/2608.11555#S5.SS5)\) is an*observational*necessary condition—a feature that fails to improve conditional likelihood when correctly supplied cannot be a useful timing signal—and we position the causal treatment as future work\.

## 3\.Models

We hold the conditioning framework fixed and vary thebackboneandcomponents\.

Backbones\.\(i\)*NHP*: continuous\-time LSTM \(CTLSTM\); the hidden state decays between events, but the LSTM attenuates long\-range history\. \(ii\)*THP*: Transformer with causal self\-attention over events; captures long\-range dependence\. In our*plain THP*ablation the intensity is fully*frozen*between events: we omit the original THP’s current\-influence termα⁡\(t−tj\)/tj\\alpha\(t\-t\_\{j\}\)/t\_\{j\}\([thp](https://arxiv.org/html/2608.11555#bib.bib3)\), so published THP is only*state*\-frozen \(its intensity is linearly time\-modulated between events\) while our ablation removes between\-event time\-dependence entirely\. \(iii\)*S2P2*: a state\-space / latent\-linear\-Hawkes layer whose intensity integral is computed in*closed form*\(no Monte\-Carlo error\)\.

Components\(on the THP backbone\)\. \(a\)*Decay head*\(“\-D”\): let the hidden state decay between events,h⁡\(t\)=h⁡\(ti\)​exp⁡\(−δ​Δ​t\)h\(t\)=h\(t\_\{i\}\)\\exp\(\-\\delta\\,\\Delta t\)with a small learnedδ\\delta; this relaxes THP’s frozen\-intensity assumption\. \(b\)*LTV*gate/shift and \(c\)*category\-aware decay*: condition the intensity on a per\-customer value signal and event category\. \(d\)*RFM*: a causal recency/frequency/cadence feature—frequency=log⁡\(1\+position\)=\\log\(1\+\\text\{position\}\)and cadence=log⁡1​p=\\log\\mathrm\{1p\}of the causal prefix\-mean of inter\-event times \(using only events up to positionii, to avoid leaking the target\)\. \(e\)*Exogenous*: per\-event season and region embeddings\. \(f\)*Continuous covariate*\(“\-W”\): a learned linear projection that injects a per\-event continuous covariate vector into the event representation \(we call it the*weather projection*after its original use case\); it carries the continuous positive control and the taxi hour\-of\-day covariate \(§[5\.5](https://arxiv.org/html/2608.11555#S5.SS5)\)\. We make*no*architectural\-novelty claim; these are standard components whose*data\-dependent value*we measure\.

## 4\.Experimental Setup

Datasets\.Three public benchmarks—Amazon \(product reviews, essentially one event type\), Taobao \(e\-commerce actions, 17 types\), RetailRocket \(very long web\-browsing sequences\)—andThumbtack, a U\.S\. home\-services marketplace where a customer sends a*paid request*to professionals for a job \(we keep customers with≥3\\geq\\\!3such requests; a customer’s per\-event “region” is their Designated Market Area, DMA—one of∼\\sim210 U\.S\. media\-market regions\)\. Public benchmarks use the EasyTPP\-Gatech splits; Thumbtack is proprietary and reported in*relative*terms\.

Metrics\.Our primary metric is thetemporal NLL\(↓\\downarrow\)\. For an event\-time sequencet1<⋯<tnt\_\{1\}<\\dots<t\_\{n\}with conditional intensityλ∗​\(t\)=λ⁡\(t∣ℋt\)\\lambda^\{\*\}\(t\)=\\lambda\(t\\mid\\mathcal\{H\}\_\{t\}\), the temporal log\-likelihood is

\(1\)log⁡L=∑i=1nlog⁡λ∗​\(ti\)−∫t0tnλ∗​\(u\)​𝑑u,\\log L=\\sum\_\{i=1\}^\{n\}\\log\\lambda^\{\*\}\(t\_\{i\}\)\\;\-\\;\\int\_\{t\_\{0\}\}^\{t\_\{n\}\}\\lambda^\{\*\}\(u\)\\,du,and we report the*per\-event*temporal NLL=−1n​log⁡L=\-\\tfrac\{1\}\{n\}\\log L\. It isolates return\-*timing*, is computed by one consistent pipeline across all datasets and model families \(the integral in closed form for S2P2\-F, by Monte\-Carlo quadrature otherwise\), and excludes the mark term\. \(Cross\-model NLL therefore mixes a closed\-form compensator for S2P2\-F with annmc=10n\_\{\\mathrm\{mc\}\}\\\!=\\\!10Monte\-Carlo estimate for the others; we treat the S2P2\-F\-vs\-MC NLL gap cautiously and rest our claims on the*same\-estimator*decay contrast \(THP vs\. THP\-D, an identical Monte\-Carlo compensator\)\.\) We also reportinter\-event RMSE/MAE\(↓\\downarrow\); these are comparable*within a dataset only*, as the time unit differs across datasets\. The same training protocol \(Adam, warmup, early stopping on validation NLL\) is used for all models, with mean±\\pmstd over seeds\{42,123,456\}\\\{42,123,456\\\}\. Throughout, we read a conditioning delta as*null*when\|Δ\|<2×\|\\Delta\|<2\\timesits seed std and as real otherwise; the taxi recovery of §[5\.5](https://arxiv.org/html/2608.11555#S5.SS5)\(≈3×\\approx\\\!3\\times\) clears this bar, and no decay\-regime conditioning delta does\.

Backbone validation\.Before drawing component conclusions we confirm our backbones reproduce known behavior: the relative ranking of S2P2, NHP, THP, and RMTPP matches the S2P2 paper\([s2p2](https://arxiv.org/html/2608.11555#bib.bib9)\)\. \(Three further reimplementations did not reproduce their published likelihoods in our pipeline; we omit them rather than report numbers we cannot trust, with details in App\.[A](https://arxiv.org/html/2608.11555#A1)\.\)

## 5\.Results

We present the evidence bottom\-up\. The decay gain \(§[5\.1](https://arxiv.org/html/2608.11555#S5.SS1)\) and backbone comparison \(§[5\.2](https://arxiv.org/html/2608.11555#S5.SS2)\) establish that the inter\-event clock dominates; the conditioning redundancy \(§[5\.3](https://arxiv.org/html/2608.11555#S5.SS3)\) and themodel\-free ceiling\(§[5\.4](https://arxiv.org/html/2608.11555#S5.SS4), contribution C2\) measure how little signal remains; and thescreen\-and\-confirm protocol\(§[5\.5](https://arxiv.org/html/2608.11555#S5.SS5), contribution C1\) is the method that makes the negatives credible\. The two framed contributions thus arrive last by design—the earlier subsections build to them\.

### 5\.1\.Decay is the dominant timing lever \(a known mechanism, quantified here\)

Adding the decay head to a plain Transformer\-Hawkes \(THP→\\toTHP\-D\) yields large temporal\-NLL gains on*every*dataset \(Table[1](https://arxiv.org/html/2608.11555#S5.T1)\): Amazon→−2\.550\.32\\\!\\to\\\!\-2\.55\(Δ−2\.9\\Delta\\,\{\-\}2\.9\), RetailRocket→−3\.530\.52\\\!\\to\\\!\-3\.53\(Δ−4\.0\\Delta\\,\{\-\}4\.0\), Taobao−→−2\.58\-2\.09\\\!\\to\\\!\-2\.58\(Δ−0\.5\\Delta\\,\{\-\}0\.5\); all mean over 3 seeds, std≤0\.09\\leq 0\.09\(Table[2](https://arxiv.org/html/2608.11555#S5.T2)\); and on Thumbtack aΔ\\Deltaof−1\.26\-1\.26\. This gain is*distributional*—it sharpens the timing*likelihood*, not point error: with the point\-prediction leak removed \(§[7](https://arxiv.org/html/2608.11555#S7)\), THP\-D’s RMSE/MAE match plain THP on every dataset \(§[5\.4](https://arxiv.org/html/2608.11555#S5.SS4)\)\. An earlier large MAE reduction we reported here was that read\-out artifact and is retracted; the point\-prediction view is the model\-free ceiling of §[5\.4](https://arxiv.org/html/2608.11555#S5.SS4)\. The gain’s*size*varies \(small on Taobao, large elsewhere\), but its*sign*is negative on every dataset—this is the most consistent result in the study\. Mechanism: decay relaxes THP’s frozen\-between\-events intensity so “time since last event” drives the rate—the classical mechanism of Hawkes processes and the neural Hawkes process\([hawkes](https://arxiv.org/html/2608.11555#bib.bib18);[nhp](https://arxiv.org/html/2608.11555#bib.bib2)\)\.*Is the gain just decay, or THP?*Continuous\-time models with built\-in decay \(NHP, and the state\-space S2P2\) likewise sit far below plain THP; the lever is the decay mechanism, not the backbone it is attached to\. \(Plain THP is a fair*ablation*—identical training and tuning, the between\-event decay the single controlled variable—not a model one would deploy; modern attention TPPs are not frozen, and our plain THP omits the original’s current\-influence term \(§3\), so the contrast quantifies the full between\-event time\-dependence channel rather than a deficit of published THP\.\) The same mechanism underlies S2P2’s strong likelihood in §[5\.2](https://arxiv.org/html/2608.11555#S5.SS2)—exponential decay is the dominant mode of its state\-space relaxation dynamics—so “decay is the lever” spans architectures, not merely the THP head\.

Table 1\.Decay gain:Δ\\Deltatemporal NLL from adding the decay head \(THP→\\toTHP\-D\); negative==better\. Mean over 3 seeds; Thumbtack relative\. Negative on every dataset—size varies \(smallest on Taobao\), sign does not\.
### 5\.2\.No single best backbone—data decides on likelihood

This comparison is*not*itself a contribution; it plays two supporting roles—confirming our reproductions track the literature \(Backbone validation, above\) and locating the decay mechanism\. Table[2](https://arxiv.org/html/2608.11555#S5.T2)reports temporal NLL \(mean±\\pmstd over three seeds\) for the models we could seed cleanly\. The robust,*same\-estimator*comparison is thedecay contrast\(THP→\\toTHP\-D, both Monte\-Carlo\): decay lowers NLL by whole nats on every dataset \(std≤0\.09\\leq 0\.09\), confirming §[5\.1](https://arxiv.org/html/2608.11555#S5.SS1)with error bars\. The closed\-form state\-space S2P2\-F is competitive\-to\-best on the public sets \(−3\.44\-3\.44Taobao,−4\.95\-4\.95RR\), but its compensator is*exact*while the THP models’ is Monte\-Carlo, so we parenthesize it as a reference and do*not*rank it against the MC models \(its Thumbtack cell is unavailable\)\. The Monte\-Carlo S2P2 and NHP variants are omitted; single runs do not change the picture\. We therefore make*no*universal\-backbone\-winner claim: decay is the*first\-order*lever \(whole nats\), residual cross\-backbone differences are*second\-order*and data\-dependent, and “sufficient” means decay reaches near the attainable ceiling—not that the backbone is irrelevant\. This is consistent with the large\-scale finding of Bosser and Ben Taieb\([bosser](https://arxiv.org/html/2608.11555#bib.bib21)\)that time\-NLL differences across neural TPP architectures are small; our multi\-nat contrasts arise from the deliberately fully\-frozen ablation \(§3\), not from cross\-architecture spread\. We also make no claim against IntensityFree/FullyNN/AttNHP \(dropped; competitive in their own papers\)\. \(Apparent point\-timing RMSE differences are largely a*read\-out*artifact, not a backbone property; §[7](https://arxiv.org/html/2608.11555#S7)\.\)

Table 2\.Temporal NLL \(↓\\downarrow\), mean±\\pmstd over 3 seeds\. Public absolute; Thumbtack relative to THP\-D \(Δ\\Delta\)\. The same\-estimator*decay contrast*\(THP vs\. THP\-D, both Monte\-Carlo\) is the robust comparison\.†S2P2\-F uses an exact closed\-form compensator \(parenthesized;*not*ranked against the MC models; Thumbtack cell unavailable\)\. Monte\-Carlo S2P2 and NHP are omitted\.*Pre\-specified sanity filter\.*We apply one fixed exclusion rule, decided from data/model properties rather than the ranking we wished to see:implausible\-given\-inputsvariants are dropped—e\.g\. browse\-augmented models reporting NLL−8\-8to−16\-16*including on Amazon, which contains no browse data*, so the “gain” cannot come from the claimed signal\.

### 5\.3\.Conditioning is redundant under decay

Why test this at all? Because adding these signals is exactly what practitioners and the recent TPP literature do \(§[2](https://arxiv.org/html/2608.11555#S2)covers covariate\-, marked\-, and external\-covariate\-conditioned models\); the screen and ceiling let us*certify*whether it actually pays for return\-timing\. We add each such signal on top of the backbone \(Table[3](https://arxiv.org/html/2608.11555#S5.T3)\)\. A*weak*\(no\-decay\) baseline does benefit from category \(CatOnlyΔ−0\.73\\Delta\\,\{\-\}0\.73\)\. Butonce the decay head is present, every conditioning signal is null or harmful: category worsens it \(CatOnly\-D\+0\.47\+0\.47; ExogCat\-D\+0\.65\+0\.65, the worst\), exogenous calendar/geography is null \(ExogTHP\-D\+0\.012\+0\.012\), RFM collapses to≈0\\approx\\\!0once its target leak is removed \(§[7](https://arxiv.org/html/2608.11555#S7)\), and proxy\-LTV does not improve likelihood \(LTV\-CTPP\-D\+0\.47\+0\.47\)\.The same holds on the public benchmarks\(Table[4](https://arxiv.org/html/2608.11555#S5.T4)\): under decay every signal moves temporal NLL by≲0\.06\\lesssim\\\!0\.06, while the weak \(no\-decay\) baseline still gains from category\. We distinguish two regimes the table conflates: on public data the decay\-regime effect is statistically*null*\(redundancy—the signal carries no information the clock lacks\); on Thumbtack adding category under decay is mildly*harmful*\(\+0\.47\+0\.47\), which we read as extra capacity overfitting in the absence of signal rather than as negative information \(the harmful deltas’ large seed variance,±0\.11\\pm 0\.11–0\.170\.17in Table[3](https://arxiv.org/html/2608.11555#S5.T3), supports the instability reading\)\. We therefore treat LTV/category/RFM/exogenous here as*measured components, not contributions*: under the signals available offline they are redundant with decay\. \(Whether*real*customer\-revenue LTV helps remains open—the offline proxy ceiling is low by construction\.\) To be explicit about scope: this is a claim about the per\-customer*timing likelihood conditional on the event history*—not about targeting, churn/propensity, mark/volume, or value models, where these same features remain informative \(§[5\.5](https://arxiv.org/html/2608.11555#S5.SS5)\)\.

Table 3\.ConditioningΔ\\Deltatemporal NLL relative to the same backbone \(↓\\downarrowbetter;\+\+is worse\), Thumbtack \(relative\); mean over 3 seeds,±\\pmthe seed std of theΔ\\Delta\. The harmful decay\-regime deltas carry large seed variance, consistent with the instability/overfitting reading \(§[5\.3](https://arxiv.org/html/2608.11555#S5.SS3)\); ExogTHP\-D is within seed noise\. Conditioning helps only the weak baseline; under decay it is null or harmful\.VariantSignal addedΔ\\Deltatemporal NLL*over THP \(no decay\)*CatOnlycategory−0\.73±0\.01\\mathbf\{\-0\.73\}\\pm 0\.01ExogTHPseason\+\+region\+0\.007±0\.005\+0\.007\\pm 0\.005*over THP\-D \(with decay\)*CatOnly\-Dcategory\+0\.47±0\.11\+0\.47\\pm 0\.11RFM\-Drecency/freq/cadence−0\.01±0\.01\-0\.01\\pm 0\.01LTV\-CTPP\-Dproxy LTV\+0\.47±0\.10\+0\.47\\pm 0\.10ExogTHP\-Dseason\+\+region\+0\.012±0\.016\+0\.012\\pm 0\.016ExogCat\-Dseason\+\+region\+\+cat\+0\.65±0\.17\+0\.65\\pm 0\.17Table 4\.ConditioningΔ\\Deltatemporal NLL on the*public*benchmarks, confirming the Thumbtack pattern: under decay every signal is≲0\.06\\lesssim\\\!0\.06; the weak \(no\-decay\) baseline still gains from category\. Mean±\\pmseed\-std ofΔ\\Deltaover 3 seeds; every decay\-regime delta is within∼1​σ\\sim\\\!1\\sigmaof zero \(null by the §4 rule\)\.
### 5\.4\.A model\-free ceiling: returns are near\-memoryless

Why does conditioning fail? Because, model\-free, there is almost nothing to condition*on*\. By a*model\-free ceiling*we mean an upper bound on predictability computed*without*fitting any TPP—the fraction of gap variance a covariate explains in a simple regression—so it bounds the*mean\-shift \(point\-prediction\)*signal available to any model\. By the reconciliation argument below it does*not*bound distributional \(likelihood\) gains; the NLL\-space redundancy claim rests on the measured deltas \(Tables[3](https://arxiv.org/html/2608.11555#S5.T3),[4](https://arxiv.org/html/2608.11555#S5.T4)\) and the screen \(§[5\.5](https://arxiv.org/html/2608.11555#S5.SS5)\), with the ceiling explaining why*point\-prediction*gains, specifically, are unavailable\. Table[5](https://arxiv.org/html/2608.11555#S5.T5)shows that no available covariate explains more than a single\-digit percentage of gap variance on*any*of the four datasets \(we report the fraction of*log*\-gap variance explained:r2r^\{2\}for the continuous previous gap, one\-wayη2\\eta^\{2\}for categoricals\): on Thumbtack the previous gap is the best single predictor at≈2\.2%\\approx\\\!2\.2\\%\(r=0\.147r=0\.147\), category≈1\.7%\\approx\\\!1\.7\\%, and season and DMA region explain effectively nothing \(η2≈5​–​8×10−4\\eta^\{2\}\\approx 5\\text\{\-\-\}8\\times 10^\{\-4\}\); on the public benchmarks the pattern repeats \(previous gap1\.41\.4–6\.3%6\.3\\%, event\-type/markη2≤1\.1%\\eta^\{2\}\\leq 1\.1\\%\)\. Two regularities: the ceiling is single\-digit everywhere, and the best single covariate is*always the previous gap*—i\.e\. the inter\-event clock itself, sequence\-derivable and already captured by decay\. Nor is the ceiling low by construction: on a genuinely clock\-driven dataset \(NYC taxi, §[5\.5](https://arxiv.org/html/2608.11555#S5.SS5)\), the same model\-free check on the exogenous hour\-of\-day covariate yieldscorr=0\.15\\mathrm\{corr\}=0\.15\(r2≈2%r^\{2\}\\approx 2\\%\)—an order of magnitude above Thumbtack’s seasonη2\\eta^\{2\}—and the conditioning screen fires there; on customer\-return data the seasonal signal simply is not present\. Moreover, behavioral surge \(a burst of sessions or a marketing touch\)*appears to carry near\-zero lead\-time*—it coincides with the return rather than preceding it by days—so it could not be used to predict a return ahead of time \(an informal observation, not a measured result; §[7](https://arxiv.org/html/2608.11555#S7)\)\. Return timing is dominated by an exogenous, near\-memoryless latent need; the inter\-event clock captured by decay is most of the recoverable signal\. Precisely, the process is near\-*renewal*: the hazard depends strongly on time since the last event \(that is what decay captures\) but on little else—“near\-memoryless” throughout means memoryless*beyond the current gap’s clock*, not a constant hazard\.

Table 5\.Model\-free fraction of log\-gap variance explained by each covariate \(r2r^\{2\}for the continuous previous gap; one\-wayη2\\eta^\{2\}for categoricals\)\. Single\-digit everywhere, and the best covariate is always the*previous gap*—the clock that decay captures\. Public columns: pooled splits, category==event type \(marks\); season/region covariates are unavailable on the public benchmarks\.Reconciling with the decay gain \(§[5\.1](https://arxiv.org/html/2608.11555#S5.SS1)\)\.A natural objection: if returns are near\-memoryless, how can decay lower NLL by22–33nats? Because the two measure different things\. The ceiling is*point\-prediction*variance—how much a covariate moves the conditional*mean*gap\. The decay gain is a*distributional*\(likelihood\) improvement: even when the mean is nearly unpredictable, modeling the*shape*of the intensity \(the post\-event refractory dip and its drift back, and the compensator∫λ∗\\int\\lambda^\{\*\}\) calibrates the timing density far better than a flat\-rate baseline\. Decay does not predict*which*gap better \(lowr2r^\{2\}\); it places probability mass over*when*far better \(largeΔ\\DeltaNLL\)\. So the findings are consistent: little is point\-predictable, yet correct temporal*calibration*is worth a lot—and it comes from the inter\-event clock, not from external covariates\. To make this concrete: on a*perfectly*memoryless process \(i\.i\.d\.Exp⁡\(λ\)\\mathrm\{Exp\}\(\\lambda\)gaps, where*no*covariate has any predictive power\), a model whose rate is off by a constant factorccpaysc−1−ln⁡cc\-1\-\\ln cexcess nats per event—1\.61\.6atc=4c\\\!=\\\!4,6\.76\.7atc=10c\\\!=\\\!10\. Multi\-nat NLL gains are thus attainable at*zero*point\-predictability: getting the rate*scale/shape*right \(what decay does, versus THP’s frozen between\-event rate\) is worth nats even when*which*gap is unpredictable\. Empirically this is exactly what we observe: once the point\-prediction leak is removed \(§[7](https://arxiv.org/html/2608.11555#S7)\), THP, THP\-D, and LTV\-CTHP\-D sit at the*same*inter\-event RMSE \(Amazon≈0\.30\\approx\\\!0\.30, Taobao≈0\.14\\approx\\\!0\.14, RetailRocket≈9\.5\\approx\\\!9\.5\), indistinguishable from predicting the global\-mean gap, and a plain RFM→\\togradient\-boosting regressor matches the neural TPPs on this point\-prediction\. Nothing beats the constant on*which*gap, while decay still wins whole nats on*when*—the point\-prediction ceiling and the distributional decay gain are two faces of the same near\-memoryless process\.

### 5\.5\.Screening exogenous features: synthetic recover, real reject

The protocol\.\(1\)Fix the candidate feature’s*encoding*and the prediction target \(here, inter\-event timing\)\.\(2\)On synthetic data, plant a feature→\\totarget coupling of tunable strengthβ\\betaand confirm the conditioned model recovers it*monotonically*inβ\\beta—a positive control that calibrates sensitivity\.\(3\)Run the*identical*pipeline on the real feature\.\(4\)Interpret*within the control’s scope*: given a passing control, a flat real result certifies “no signal in that encoding,” not a weak method\. The pattern ports standard positive/negative\-control practice—placebo and refutation tests in econometrics, sanity checks in ML\([adebayo](https://arxiv.org/html/2608.11555#bib.bib23)\)—to TPP conditioning\.

A null is only informative if the method*could*have found a signal\. We therefore validate the exogenous pipeline with apositive control: we generate synthetic sequences in which an exogenous field \(season\) is coupled to the next gap at a tunable strengthβ\\beta\(a causal coupling that a decay\-only model cannot recover, since the field is i\.i\.d\. across events\), and check that a season/region\-conditioned model recovers it\. It does,*monotonically*inβ\\beta\(Fig\.[1](https://arxiv.org/html/2608.11555#S5.F1)\): ExogTHP improves by−0\.059\-0\.059,−0\.232\-0\.232,−0\.824\-0\.824atβ=0\.5,1,2\\beta=0\.5,1,2\(and\+0\.001\+0\.001atβ=0\\beta=0, correctly null\); the decay variant ExogTHP\-D recovers as well \(−0\.042,−0\.143,−0\.337\-0\.042,\-0\.143,\-0\.337\)\. The protocol is*not*specific to categorical features: replacing the categorical season/region embedding with a learned linear projection of a per\-event*continuous*covariatezz, and planting a continuous couplinggap∼Exp⁡\(eβ​z\)\\text\{gap\}\\sim\\mathrm\{Exp\}\(e^\{\\beta z\}\)withz∼𝒩⁡\(0,1\)z\\sim\\mathcal\{N\}\(0,1\), yields the same monotonic recovery under decay—Δ\\Deltatime\-NLL\+0\.000,−0\.092,−0\.258,−0\.738\+0\.000,\-0\.092,\-0\.258,\-0\.738atβ=0,0\.5,1,2\\beta=0,0\.5,1,2\. Both instantiations are correctly null atβ=0\\beta=0\. The same pipeline, applied to real calendar/geography \(Thumbtack\), yields no improvement \(Table[3](https://arxiv.org/html/2608.11555#S5.T3)\)\.

The real null is genuine, not a plumbing artifact\.We verified that the exogenous signal reaches the model with rich variation: Thumbtack carries all 12 calendar months and 207 DMA regions \(99\.8%99\.8\\%non\-zero\), well\-distributed, with loader keys matching the data\. Because the positive control confirms “signal present⇒\\Rightarrowrecovered,” the real\-data null means “nothing to recover”—calendar and geography do not move the per\-customer conditional gap\. This screen\-and\-confirm pattern is the recommended way to read a conditioning null\.

Scope of the claim\.The screen adjudicates a*specific*pair: the candidate feature*as encoded*and the chosen target \(here, inter\-event timing\)\. A null therefore means “no signal in that encoding for that target,” bounded by what the positive control shows recoverable—*not*“no exogenous signal of any kind\.” Likewise, of the signals in C3, calendar/geography passed through this planted\-coupling screen; the LTV/category/RFM nulls are*measured*under the same pipeline and bounded by the same ceiling but were not separately screened—per\-pathway positive controls are future work\. This boundary is consistent with the broader literature: on canonical weather\-sensitive data the exogenous covariate drives*volume and marks*rather than per\-event timing\. As a model\-free check, Beijing PM2\.5 alert*onsets*are near\-uniform across calendar months and their inter\-onset gaps are essentially season\-invariant, and on the Walmart weather dataset temperature is uncorrelated with the inter\-sale gap \(r=0\.005r\\\!=\\\!0\.005\); a*timing*screen is thus correctly null there, even though a mark/volume model benefits from the same covariate\([transfeat](https://arxiv.org/html/2608.11555#bib.bib13);[metp](https://arxiv.org/html/2608.11555#bib.bib14)\)\. Extending the screen to continuous*and*lagged encodings \(each with its own positive control\) is a direct generalization\.

A real positive control\.The screen also fires on a*real*exogenous signal where one genuinely exists\. On NYC green\-taxi pickups \(Jan 2019\), inter\-pickup gaps are modulated by*hour\-of\-day*\(rate high at rush hour, low overnight; model\-freecorr⁡\(sin⁡hour,log⁡gap\)=0\.15\\mathrm\{corr\}\(\\sin\\text\{hour\},\\log\\text\{gap\}\)=0\.15\)\. Encoding hour as a continuous cyclic covariate, the screen lowers temporal NLL byΔ=−0\.025\\Delta\\\!=\\\!\-0\.025\(→1\.7311\.756\\\!\\to\\\!1\.731;n=3n\\\!=\\\!3, well outside seed noise±0\.008\\pm 0\.008; replicated on a second month, Feb 2019,Δ=−0\.021\\Delta\\\!=\\\!\-0\.021; hour is genuinely exogenous—each sequence starts at its zone\-day’s first pickup, so cumulative time does not encode the clock hour\)—a*modest*recovery that tracks the*modest*real signal, in contrast to the strong syntheticβ\\betaand the flat \(≈0\\approx\\\!0\) real calendar/geography\. This closes the loop: synthetic controls calibrate sensitivity, real negatives \(PM2\.5, Walmart\) and a real positive \(taxi\) confirm the screen discriminates on real data—so the customer\-return null is a property of that data, not a weakness of the method\. We still describe screen\-and\-confirm as a*validation discipline*\(establish recoverability, then trust the null\) rather than a turnkey discovery method\.

000\.50\.51122−0\.8\-0\.8−0\.6\-0\.6−0\.4\-0\.4−0\.2\-0\.200planted strengthβ\\betaΔ\\DeltaNLL \(↓\\downarrow\)\(a\) categorical season embeddingExogTHPExogTHP\-D \(decay\)

000\.50\.51122−0\.8\-0\.8−0\.6\-0\.6−0\.4\-0\.4−0\.2\-0\.200planted strengthβ\\betaΔ\\DeltaNLL \(↓\\downarrow\)\(b\) continuous covariate \(learned projection\)THP\-D\-W \(decay\)

Figure 1\.Positive control \(two encodings\)\.A planted exogenous→\\togap coupling of strengthβ\\betais recovered*monotonically*in both:\(a\)a*categorical*season embedding \(no\-decay ExogTHP and decay ExogTHP\-D\), and\(b\)a*continuous*covariate via a learned linear projection \(decay THP\-D\-W\)\.β=0\\beta=0is correctly null in both, so the screen is not categorical\-specific\. This is what a*real*signal would look like—real calendar/geography \(Table[3](https://arxiv.org/html/2608.11555#S5.T3)\) does not\.

## 6\.Mechanism: why decay absorbs the rest

The redundancy has a*dual mechanism*\.\(A\) Sequence\-derivable signals are already captured\.RFM and category are deterministic functions of the observed event history, which the decay backbone already encodes; supplying them explicitly adds no information \(RFM\-D≈0\\approx\\\!0, CatOnly\-D worse once decay is in\)\.\(B\) Exogenous signals do not move the conditional gap\.Season and region shift*demand composition*\(who buys what, where\) but not the per\-customer*timing*of the next return, which is driven by an exogenous latent need realized same\-day \(§[5\.4](https://arxiv.org/html/2608.11555#S5.SS4)\)\. Timing≠\\neqdemand\-composition\. Together these explain why a single inter\-event\-clock mechanism \(decay\) is sufficient and the usual enrichments are redundant\.

What this buys the practitioner\.For the notification\-timing and CRM use cases of §1: a decay\-calibrated intensity supports return\-*window*estimation \(when the hazard recovers from the post\-event refractory dip\) and send suppression during the dip; it does*not*support point\-timed sends \(point prediction sits at the global\-mean ceiling, §[5\.4](https://arxiv.org/html/2608.11555#S5.SS4)\), and the effect of*timing an intervention*is a causal question requiring interventional data \(§[2](https://arxiv.org/html/2608.11555#S2), future work\)\.

## 7\.Limitations and honest evaluation

We report results we initially got wrong and corrected, in the spirit of honest evaluation\. An NHP target\-indexing bug once produced a fake7×7\\timesMAE improvement; an RFM cadence feature once leaked the target through a full\-sequence normalizer; an S2P2 RMSE was read from an untrained auxiliary head; and the decay head’s point\-prediction read the hidden state*after*decaying it by the actual next inter\-event interval, leaking the very gap it predicts \(a counterfactual that alters only the held\-out gap moves the prediction in lock\-step,corr≈1\\mathrm\{corr\}\\\!\\approx\\\!1\)—which had inflated the decay models’ RMSE/MAE and the apparent point\-timing decay gains\. Each was caught by sanity\-checking predictions against targets; we fixed the decay read\-out to use the pre\-decay \(history\-only\) state, added a counterfactual regression test, and report only the corrected, leak\-free numbers, which show*decay\-flat*RMSE/MAE\. Temporal NLL—a density evaluated at the observed event times and treated identically across all models—is unaffected by these point\-prediction read\-out bugs\. Separately, in earlier ranking experiments a TPP’s apparent ranking ability was dominated by the score*read\-out*\(intensity integral vs\. a learned time head\) and the held\-out*anchor*rather than the architecture; an apparent “likelihood\-vs\-ranking decoupling” did not survive re\-validation, so we make*no*ranking claim\. Scope limits: one real dataset, reported only in relative terms; the customer filter \(≥3\\geq\\\!3paid requests\) conditions on repeat return, so seasonality acting on acquisition or the*first*return is excluded by construction; proxy\-LTV understates real\-LTV conditioning \(open\); and our S2P2 point\-timing is read from an auxiliary head, so we do not treat its RMSE as representative of the method\. Three claims rest on argument more than on a dedicated experiment, which we flag: the model\-free*season/region*ceilings \(§[5\.4](https://arxiv.org/html/2608.11555#S5.SS4)\) are measured on Thumbtack only \(the public benchmarks carry no calendar/geography covariates\), and platforms with strong external seasonality may sit under a higher ceiling, so the one\-day ceiling\-plus\-screen check should be re\-run per dataset instead of assuming our≲5%\\lesssim\\\!5\\%figure; the “decay already encodes RFM/category” mechanism \(§[6](https://arxiv.org/html/2608.11555#S6)\) is*inferred*, not probed \(no representation\-probing experiment\); and the “surge has zero lead\-time” observation is from exploratory analysis without a dedicated figure\.

## 8\.Conclusion

Two tools, a model\-free ceiling and a screen\-and\-confirm certification, let us verify a clean result for customer\-return*timing*: continuous\-time decay \(a known mechanism\) is nearly*sufficient*, and once it is present, the conditioning practitioners add is redundant or harmful\. A model\-free analysis shows why: returns are near\-memoryless, with≲5%\\lesssim\\\!5\\%of gap variance explainable \(and, informally, near\-zero surge lead\-time\)\. Because a conditioning null is only meaningful against a positive control, we recommend ascreen\-and\-confirmprotocol: confirm the pipeline recovers a planted synthetic signal, then trust the real\-data null\. Future work: a causal/counterfactual TPP treatment of*interventions*\(e\.g\. marketing\), which needs interventional or randomized\-holdout data; representation\-probing \(or mutual\-information\) experiments on the decay backbone’s hidden state, to turn the*inferred*“decay already encodes RFM/category” mechanism of §[6](https://arxiv.org/html/2608.11555#S6)into a measured one; and generalizing screen\-and\-confirm into automatic exogenous\-feature discovery\.

## Appendix AReproducibility details

Splits\.Public benchmarks use the EasyTPP\-Gatech\([easytpp](https://arxiv.org/html/2608.11555#bib.bib8)\)train/dev/test splits; Thumbtack uses a chronological per\-customer split, reported only in relative terms\.Optimization \(all models\)\.Adam, learning rate10−310^\{\-3\}, linear warmup over 5 epochs thenReduceLROnPlateau\(factor0\.50\.5, patience55, min lr10−510^\{\-5\}\), gradient clipping at1\.01\.0, up to5050epochs with early stopping \(patience1010\) on validation temporal NLL\. Hidden sizedmodel=64d\_\{\\text\{model\}\}=64; batch size6464\(RetailRocket1616, sequences truncated to512512\)\.±\\pmis mean/std over seeds\{42,123,456\}\\\{42,123,456\\\}\.Backbones\.THP: causal self\-attention, intensity constant between events \(the original’s current\-influence termα⁡\(t−tj\)/tj\\alpha\(t\-t\_\{j\}\)/t\_\{j\}is omitted; §3\); “\-D” adds the per\-dimension hidden\-state decay head \(small learnedδ\\delta\)\. NHP: continuous\-time LSTM\.S2P2 is our reimplementationof a continuous\-time state\-space / latent\-linear\-Hawkes layer \(dstate=64d\_\{\\text\{state\}\}=64,22layers\);*\-F*computes∫λ∗\\int\\lambda^\{\*\}in closed form \(nmc=0n\_\{\\text\{mc\}\}=0\), the non\-F variant usesnmc=10n\_\{\\text\{mc\}\}=10\. We do not use the original authors’ code\.Dropped baselines\.IntensityFree\([intfree](https://arxiv.org/html/2608.11555#bib.bib6)\), FullyNN\([fullynn](https://arxiv.org/html/2608.11555#bib.bib7)\), and AttNHP\([attnhp](https://arxiv.org/html/2608.11555#bib.bib5)\)reimplementations did not reproduce published likelihoods in our pipeline \(e\.g\. Taobao IntensityFree LL−1\.43\-1\.43vs\.\+1\.318\+1\.318reported\), so we exclude their numbers and cite the originals, which report competitive results\.Synthetic positive control\.Per\-event season∼Unif​\{1\.\.12\}\\sim\\mathrm\{Unif\}\\\{1\.\.12\\\}i\.i\.d\.;gapj\+1∼Exp⁡\(r0​exp⁡\(β​sin⁡\(2​π​seasonj/12\)\)\)\\text\{gap\}\_\{j\+1\}\\sim\\mathrm\{Exp\}\(r\_\{0\}\\exp\(\\beta\\sin\(2\\pi\\,\\text\{season\}\_\{j\}/12\)\)\)using the*source*event’s season \(causal\); region is an i\.i\.d\. noise control\. A pre\-training validation confirms the planted correlation grows withβ\\beta\. The continuous variant drawsz∼𝒩⁡\(0,1\)z\\sim\\mathcal\{N\}\(0,1\)per event and setsgapj\+1∼Exp⁡\(eβ​zj\)\\text\{gap\}\_\{j\+1\}\\sim\\mathrm\{Exp\}\(e^\{\\beta z\_\{j\}\}\), fed through the weather projection\.Model\-free ceiling \(public columns of Table[5](https://arxiv.org/html/2608.11555#S5.T5)\)\.Computed over pooled train/dev/test splits, onlog\\logof positive inter\-event gaps:r2r^\{2\}from the Pearson correlation of adjacent within\-sequence gap pairs\(log⁡gj,log⁡gj\+1\)\(\\log g\_\{j\},\\log g\_\{j\+1\}\);η2\\eta^\{2\}from a one\-way decomposition oflog⁡gj\+1\\log g\_\{j\+1\}grouped by the source event’s type \(ceiling\_public\.py\)\. Applied to the Thumbtack event data, the same script yieldsr2≈2\.0%r^\{2\}\\\!\\approx\\\!2\.0\\%and categoryη2≈1\.1%\\eta^\{2\}\\\!\\approx\\\!1\.1\\%, consistent in magnitude with the cohort\-derived Thumbtack column of Table[5](https://arxiv.org/html/2608.11555#S5.T5)\.Real positive control \(taxi\)\.NYC green\-taxi pickups, Jan 2019; sequences are per \(pickup\-zone, day\) with≥20\\geq\\\!20pickups \(capped at 200,n=4000n\\\!=\\\!4000\); the per\-event covariate is hour\-of\-day encoded as\[sin⁡\(2​π​h/24\),cos⁡\(2​π​h/24\),0\]\[\\sin\(2\\pi h/24\),\\cos\(2\\pi h/24\),0\]through the same weather projection\. Events are single\-type \(timing only\)\.

## References

- \(1\)N\. Du et al\. Recurrent Marked Temporal Point Processes\.*KDD*, 2016\.
- \(2\)H\. Mei and J\. Eisner\. The Neural Hawkes Process\.*NeurIPS*, 2017\.
- \(3\)S\. Zuo et al\. Transformer Hawkes Process\.*ICML*, 2020\.
- \(4\)Q\. Zhang et al\. Self\-Attentive Hawkes Process\.*ICML*, 2020\.
- \(5\)C\. Yang, H\. Mei, J\. Eisner\. Transformer Embeddings of Irregularly Spaced Events and Their Participants \(AttNHP\)\.*ICLR*, 2022\.
- \(6\)O\. Shchur et al\. Intensity\-Free Learning of Temporal Point Processes\.*ICLR*, 2020\.
- \(7\)T\. Omi, N\. Ueda, and K\. Aihara\. Fully Neural Network based Model for General Temporal Point Processes\.*NeurIPS*, 2019\.
- \(8\)S\. Xue et al\. EasyTPP: Towards Open Benchmarking Temporal Point Processes\.*ICLR*, 2024\.
- \(9\)Y\. Chang, A\. Boyd, C\. Xiao, T\. Kass\-Hout, P\. Bhatia, P\. Smyth, and A\. Warrington\. Deep Continuous\-Time State\-Space Models for Marked Event Sequences\.*NeurIPS*, 2025\.
- \(10\)O\. Shchur et al\. Neural Temporal Point Processes: A Review\.*IJCAI*, 2021\.
- \(11\)G\. Grob et al\. A Recurrent Neural Network Survival Model: Predicting Web User Return Time\.*ECML PKDD*, 2018\.
- \(12\)K\. Noorbakhsh and M\. Gomez\-Rodriguez\. Counterfactual Temporal Point Processes\.*NeurIPS*, 2022\.
- \(13\)Z\. Meng, B\. Li, X\. Fan, Z\. Li, Y\. Wang, F\. Chen, and F\. Zhou\. TransFeat\-TPP: An Interpretable Deep Covariate Temporal Point Processes\.*ECAI*, 2024\.
- \(14\)B\. Li, L\. Zhang, F\. Tsung, and X\. Zhang\. METP: Multi\-Granularity Integration of External Covariates for Temporal Point Processes\.*AAAI*, 2026\.
- \(15\)P\. Fader, B\. Hardie, K\. Lee\. “Counting Your Customers” the Easy Way: BG/NBD\.*Marketing Science*, 2005\.
- \(16\)D\. Schmittlein, D\. Morrison, R\. Colombo\. Counting Your Customers: Pareto/NBD\.*Management Science*, 1987\.
- \(17\)X\. Wang, T\. Liu, J\. Miao\. A Deep Probabilistic Model for Customer Lifetime Value \(ZILN\)\. arXiv:1912\.07753, 2019\.
- \(18\)A\. G\. Hawkes\. Spectra of Some Self\-Exciting and Mutually Exciting Point Processes\.*Biometrika*, 1971\.
- \(19\)K\. Kapoor, M\. Sun, J\. Srivastava, and T\. Ye\. A Hazard Based Approach to User Return Time Prediction\.*KDD*, 2014\.
- \(20\)N\. Du, Y\. Wang, N\. He, J\. Sun, and L\. Song\. Time\-Sensitive Recommendation from Recurrent User Activities\.*NeurIPS*, 2015\.
- \(21\)T\. Bosser and S\. Ben Taieb\. On the Predictive Accuracy of Neural Temporal Point Process Models for Continuous\-time Event Data\.*TMLR*, 2023\.
- \(22\)M\. Platzer and T\. Reutterer\. Ticking Away the Moments: Timing Regularity Helps to Better Predict Customer Activity\.*Marketing Science*, 2016\.
- \(23\)J\. Adebayo, J\. Gilmer, M\. Muelly, I\. Goodfellow, M\. Hardt, and B\. Kim\. Sanity Checks for Saliency Maps\.*NeurIPS*, 2018\.

相似文章

改进的幻象:信用评分中的拒绝推断策略

arXiv cs.LG

本文系统评估了信用评分中的拒绝推断方法,并发现了一种结构性失效模式:在自然的再训练周期中,模型的准确率提升但召回率骤降,造成了改进的幻象,而实际拒绝质量却在恶化。本文提出了一种受控探索策略,无需统计假设即可打破反馈循环,并证明即使最低的探索率也足以诊断该问题。