Lot 003

Treverse Recsys: temporal recommendations from learning to serving

This is a condensed version of Xuhui’s full write-up, which covers the model, the release policy, a worked cost case study and the experiment in much more depth. Read the complete article on xuhuizhan5.github.io.

What should a recommender learn when an item may leave the auction before its behavioral history becomes useful?

That question shaped Treverse Recsys, the recommendation system behind Haggle. Newly eligible items arrive cold, older items disappear from the candidate set, and user histories range from rich to nearly absent. The clock changes both what the model can know and what it is still allowed to recommend.

The system answers by learning from time-safe relational evidence, doing the expensive ranking work offline, and publishing a verified decision artifact. The online service then has a smaller job: load a complete release, choose the correct recommendation path, record exposure, and fail predictably.

Cold start is the normal case

An ID-only recommender gets stronger when the same users and items repeat. In an auction that rarely happens. Fresh items arrive before they can build much behavior, older items leave quickly, and a completed outcome may settle long after the shallow action that preceded it. That produced three requirements that had to hold together:

  • The model must represent unseen inventory from context rather than waiting for an item ID to become warm.
  • Evaluation must preserve time direction so future interactions cannot leak into earlier decisions.
  • Online serving must keep a predictable request path without making neural inference another request-time dependency.

The product objective mattered just as much. An item open is useful diagnostic evidence, but it is not interchangeable with a bid, a checkout, or realized marketplace value. Optimizing the easiest action would have produced a cleaner dashboard and potentially a weaker business result.

A temporal graph for the shape of the problem

Popularity, matrix factorization, two-tower retrieval and sequential models were all considered, and several remain in use as baselines, benchmarks or fallbacks. The chosen retrieval family is a temporal graph, because a bid is a timestamped relationship between a participant and an evolving item, not merely metadata or a generic click.

The retriever is GraphSAGE-style rather than a literal reproduction of one paper: learn an aggregation function instead of one embedding per permanent node. It combines content encoders with edge-aware neighborhood aggregation, filters sampled history by the query time, and represents recency and interaction strength in edge context. A learned gate can lean toward content when graph evidence is weak, which matters for both new inventory and sparse users. Training uses contrastive retrieval with in-batch negatives plus a hard-negative term. Exact filtered retrieval within the eligible candidate scope remains the correctness reference; approximate search is an optimization for later, after a measured need.

Move intelligence off the hot path

A recommendation request should consume a decision, not recreate the learning system. Databricks curates and validates upstream. Scheduled orchestration builds features, trains or refreshes a candidate, evaluates the eligible release, and materializes ranked output. The serving process consumes the immutable ranked artifact. Live events inform measurement and later learning, but they never mutate the active model during a request, and a failed training job cannot partially update serving state. The orchestration, deployment-state and observability contracts underneath all of this are FORGE, covered in Lot 004.

A model release is also a data release. The served decision depends on features, eligibility and ranked output, not just weights, so release identity binds all of them together. The same eligibility definition applies to champion and challenger; otherwise an apparent gain may only be a candidate-set mismatch. Most releases refresh ranks without retraining. A challenger has to pass quality, coverage and scope gates, and failure keeps the champion or the last-known-good release.

The smallest complete decision service

The service resolves a versioned release, streams the snapshot into replacement maps, validates them, and derives compact serving indexes, none of it visible to requests while it is being assembled. Publication happens inside one write-locked critical section, so readers observe the old complete state or the new complete state, never a partial mixture. If refresh fails after a valid artifact has loaded, readiness stays true and health reports stale state. A deterministic fallback covers missing personalized state, and its use is observable rather than silent.

Keeping the state local in memory, duplicated across two replicas, was also the cheaper option. A worked example at 6.5 GB of state and a 48.23 QPS single-replica failover gate came to $238.27 a month on two larger workers, against $517.74 with Valkey and $712.29 with Redis OSS on smaller workers. The break-even is small, around 1.94 GB for Valkey and 1.31 GB for Redis OSS, and the migration triggers are written down in advance rather than discovered under load.

Fewer opens, stronger downstream outcomes

The production experiment used server-side assignment with committed exposure. Effects are CUPED-adjusted, with 95% confidence intervals and Holm correction across the five-metric family.

MetricRelative change95% CI
Marketplace sales per exposed buyer+12.62%+5.67% to +20.03%
Checkout-user rate+3.21%+0.65% to +5.83%
Bid-action users+2.13%+0.57% to +3.72%
Item-open users−4.56%−7.64% to −1.39%

Downstream value went up while opens went down. The reading is that the system removed low-value curiosity clicks and helped buyers reach actionable inventory more efficiently. Random assignment supports the comparison between arms; it does not identify the mechanism, and every later change, from the cold-start policy to the reranker to an intent-history challenger, gets its own controlled test.

What comes next

The roadmap is a sequence of hypotheses with triggers, not a complexity commitment. Approximate nearest-neighbor search, partitioning and managed state are all deferred until a specific bottleneck is measured. The standard for every step is the same: correctness must be replayable, availability testable, cost recomputable, and the product effect measurable.

The full write-up, including the architecture figures, the release policy, the complete cost model and the experiment forest plot, is at xuhuizhan5.github.io/project/treverse-recsys.