AI in investment · Editorial

What belongs in an AI signal's paper trail?

A confident output deserves a confident receipt. The documentation an AI investment signal should carry, item by item.

Investment signals arrived from machine learning wearing their best suit: a number, a direction, a probability. What they rarely arrive with is a receipt. Yet every machine-derived signal is an act of data assembly — and assembly leaves a trail. When that trail is missing, you are not holding research; you are holding a claim with good posture.

This desk treats the dossier as part of the signal. A model whose documentation cannot answer basic questions about its own construction is not more mysterious, only less finished. Below is the receipt we ask for, section by section, before writing about any investment signal.

The universe, exactly

A signal speaks about a universe — the set of assets it claims to cover — and the universe is where most quiet fibs live. Is membership defined at the start of the period (as it would be in live trading) or by looking across time? Does the universe include delisted securities, or only the survivors? Are the boundaries mechanical — a rule — or discretionary, adjusted by hand when results displease?

A cleanly stated universe rule sounds like legislation: precise, testable, slightly dull. A vague one uses words like “major,” “liquid” and “relevant.” When a vendor cannot restate the universe as an algorithm, the honest conclusion is that the rule is the results.

Where the features came from

Every AI signal digests features — input variables built from data — and each feature has its own lineage. The questions are mundane and decisive: which raw feeds did the features come from, who licenses them, what corrections were applied before the model saw them, and how were entities matched across sources? A model fed on features that quietly changed definition in mid-history is not one model; it is two models wearing the same name, and only one of them earned the backtest.

What the model saw while training

Training documentation answers the question when did the model learn what it knows. A window that spans only calm markets, or only one interest-rate era, produces a signal that mistakes its childhood for the world. The disclosure should state the training period plainly, distinguish it from the test period, and describe what happened to the signal when the world moved outside its experience. If the answer is silence, read the silence: the signal has never been outside the lab.

Turnover and capacity, the unglamorous truth

Two numbers describe what it costs to follow a signal: turnover — how often holdings change — and capacity — how much money can follow it before its own weight spoils the effect. Neither flatters a pitch deck, which is precisely why they belong in the receipt. A signal that only survives at breakfast-table scale is a research curiosity; a signal that needs daily turnover is a trading-cost story first and an intelligence story second.

The tells of a thin dossier

When reading a signal’s documentation, this desk watches for a short list of tells:

  • The universe is described with adjectives instead of rules.
  • Feature sources are named but licenses and revision policies are not.
  • Training and test periods are not separated, or overlap quietly.
  • Turnover and capacity appear nowhere, as if following the signal were free.
  • Performance is the only section longer than a paragraph.

Any one tell is a caution. Several together mean the dossier is decoration.

An endnote on what this page is

This essay is descriptive, not advisory: it explains what documentation contains, never whether any particular signal deserves your money — this desk sells nothing, recommends nothing, and knows nothing about your situation. If a dossier you are reading raises questions the vendor cannot answer, the inquiry form reaches the desk, and questions worth a full essay usually become one.