Explainer
Bad ticks and stale prices: how to detect them in market data
Most bad data is not dramatic. It is a price that stopped moving an hour ago and still looks perfectly reasonable on a chart.
On this page
Key takeaways
- A bad tick is a price or quote that does not reflect the tradable market: a spike, a zero, a crossed quote, a unit or decimal error, or a value that stopped updating.
- Stale data is the hardest case because every field looks valid; only the timestamp, judged against a session-aware freshness budget, gives it away.
- Silence is normal for a stock after the close and for a bond yield between daily observations, so freshness budgets must depend on asset class and session.
- For spikes, use a rolling median and median absolute deviation rather than a mean and standard deviation, and quarantine suspect ticks instead of deleting them.
- Cross-checks between trade and quote, REST and stream, and related instruments catch errors that no single series can reveal.
A bad tick is a price or quote that does not represent the tradable market: a print ten percent away from its neighbours, a zero bid, a bid above the ask, a price in cents where you expected dollars, or a stale value that quietly stopped updating. Catch them in layers: structural checks, session-aware freshness budgets, a spike filter on robust statistics, and cross-checks against related data.
No feed is immune, which is why these checks belong in your code as well as upstream. This explainer describes the approach and shows tested Python. It deliberately publishes no one's internal thresholds: the right numbers depend on your instruments and your tolerance, and the code learns them from your own history. For how we measure our own feed, see data quality.
A field guide to bad ticks
| Type | What it looks like | Typical cause | Caught by |
|---|---|---|---|
| Spike | One print far from its neighbours, then back | Erroneous trade, parse error, odd print | Spike filter |
| Zero or negative | bid: 0 or price: 0 | Empty book side, a missing value defaulted to zero | Structural check |
| Crossed or locked | Bid at or above the ask | The two sides updated at slightly different moments | Structural check |
| Unit or scale error | A price 100 times too large or small | Cents against dollars, a shifted decimal | Unit table, ratio to previous close |
| Stale or frozen | Same value, old timestamp | Disconnected source, stuck cache, closed market | Freshness budget |
| Out of order | A timestamp earlier than the last one | Late delivery, a replay after reconnect | Monotonic check per symbol |
| Wrong time unit | Dates in 1970 or far in the future | Seconds read as milliseconds, or the reverse | Plausible-range check |
| Duplicate | The same trade twice | Replay after a reconnect | Dedupe on symbol, time, price and size |
Unit errors deserve a concrete case. SUGARUSD is quoted in US cents per pound, so a snapshot bid of 18.537 means 18.5 cents, not $18.54. A pipeline that assumes dollars will either flag every sugar tick as a 99% crash or, worse, store it. Keep a unit table per instrument; the sugar price page shows the documented unit. Time units fail the same way: a Unix timestamp in seconds read as milliseconds lands in January 1970, and the Unix timestamp guide shows how to tell them apart.
Layer 1: structural checks on every message
Start with checks that need no history: every numeric field parses (WebSocket frames for crypto, forex and stocks send numbers as strings), both sides are positive, the ask is above the bid, the timestamp is in a plausible millisecond range, and it is not earlier than the last one you accepted for that symbol. These catch zeros, crossed quotes, unit mix-ups in time and out-of-order replays for almost no CPU. The bid-ask spread explainer covers why crossed and locked quotes happen.
Layer 2: stale data and freshness budgets per asset and session
Stale data is the dangerous case because nothing about the value looks wrong. The only signal is time: how long since this symbol last updated, compared with how long it normally stays quiet at this hour. That comparison needs two inputs, a budget per symbol and session, and the market status, because outside the session silence is expected and the right output is "closed, last value at 15:30", not an alarm.
| Asset | When silence is normal | How to set the budget |
|---|---|---|
| Crypto | Never fully; nights and weekends are thinner | From your own quiet-hour gaps, per pair |
| Forex | The weekend close; thin minutes around the daily rollover | Short on weekdays, suspended at the weekend |
| US stocks | Outside 04:00 to 20:00 New York, holidays, halts | Per symbol and per session; extended hours need a wider budget |
| Commodities | The daily 17:00 to 18:00 New York pause and the weekend | Session-aware, like forex |
| Indices | Outside the underlying market's session | Follow that market's calendar |
| Bond yields | Between daily observations | In days: a Friday observation read on Monday is current |
| 1-minute bars | A minute with no trades produces no bar | Never alarm on a missing bar for a thin symbol |
When silence is expected, weekday in UTC
Hours in UTC
Do not hard-code budgets from a blog post, including this one. Record the gaps between updates for each symbol during normal sessions, take a high percentile, multiply by a safety factor, and review the rejections. A liquid pair and a thin stock can differ by orders of magnitude, and so can the same stock in regular and extended hours.
Layer 3: spike filters and outlier detection in a time series
The textbook z-score, distance from the mean in standard deviations, fails exactly when you need it: a large spike inflates the standard deviation it is measured against. Use the median and the median absolute deviation (MAD) instead. They barely move when one value is wild, so the spike stands out.
robust z = |x − median(window)| ÷ (1.4826 × MAD(window))
- window
- The last N accepted values for this symbol, for example mids.
- MAD
- Median of absolute deviations from the median.
- 1.4826
- Scales MAD to match a standard deviation on normally distributed data.
A one-tick spike against the rolling median
- Price
- Rolling median
The second trick is confirmation. A bad tick reverts; a real move persists. So hold a suspect value for one tick: if the next tick lands near the suspect, accept both and rebuild the window at the new level; if it lands back near the median, drop the suspect. You pay one tick of delay on genuine gaps, which is the right trade for anything that alerts or acts.
from __future__ import annotations
import statistics
import time
from collections import deque
LAST_TS: dict[tuple[str, str], int] = {}
def structural_problem(msg: dict, kind: str = "quote") -> str | None:
"""Why a quote or trade frame is unusable, or None if it passes."""
try:
ts = int(msg.get("ts") or msg["timestamp"])
if kind == "quote":
bid, ask = float(msg["bid"]), float(msg["ask"]) # WebSocket numerics can be strings
else:
price = float(msg["price"])
except (KeyError, TypeError, ValueError):
return "missing or non-numeric field"
if not 946_684_800_000 <= ts <= time.time() * 1000 + 5_000:
return "timestamp out of range (seconds instead of milliseconds?)"
key = (msg["symbol"], kind)
if ts < LAST_TS.get(key, 0):
return "out of order"
if kind == "quote" and (bid <= 0 or ask <= 0):
return "zero or negative side"
if kind == "quote" and ask <= bid:
return "crossed or locked"
if kind == "trade" and price <= 0:
return "zero or negative price"
LAST_TS[key] = ts
return None
def learn_budget(gaps_s: list[float], floor_s: float = 1.0) -> float:
"""Freshness budget from your own history: 3x the 99th percentile gap."""
p99 = statistics.quantiles(gaps_s, n=100)[98]
return max(floor_s, 3 * p99)
def freshness(ts_ms: int, budget_s: float, status: str) -> str:
"""Classify the last value of one symbol as live, stale or closed."""
if status == "closed":
return "closed" # expected silence: show the last value and its time
age_s = time.time() - ts_ms / 1000
return "stale" if age_s > budget_s else "live"
class SpikeFilter:
"""Rolling median and MAD with one-tick confirmation."""
def __init__(self, window: int = 50, z_max: float = 8.0, min_rel_scale: float = 1e-5):
self.values: deque[float] = deque(maxlen=window)
self.z_max = z_max
self.min_rel_scale = min_rel_scale
self.pending: float | None = None
def _z(self, x: float) -> float:
med = statistics.median(self.values)
mad = statistics.median(abs(v - med) for v in self.values)
scale = max(1.4826 * mad, abs(med) * self.min_rel_scale)
return abs(x - med) / scale
def check(self, x: float) -> str:
if len(self.values) < 10:
self.values.append(x)
return "warming up"
if self.pending is not None:
suspect, self.pending = self.pending, None
if abs(x - suspect) <= abs(suspect) * 0.001: # next tick agrees with the jump
self.values.clear() # new price level: rebuild the window from here
self.values.extend([suspect, x])
return "accepted: move confirmed"
if self._z(x) > self.z_max:
self.pending = x
return f"quarantined (z = {self._z(x):.1f})"
self.values.append(x)
return "ok"import time
from tick_checks import SpikeFilter, freshness, learn_budget, structural_problem
# 1. Structural checks on stream-style frames (numerics arrive as strings)
print(structural_problem({"symbol": "EURUSD", "bid": "1.137107", "ask": "1.137157", "ts": 1790591203002}))
print(structural_problem({"symbol": "EURUSD", "bid": "1.137157", "ask": "1.137107", "ts": 1790591204002}))
print(structural_problem({"symbol": "US:KO", "price": "88.24", "ts": 1790589160}, kind="trade"))
# 2. Freshness: budget learned from observed gaps (seconds), then applied
budget = learn_budget([0.3, 0.5, 1.2, 0.8, 2.0, 0.4, 0.6, 0.9, 1.1, 3.5, 0.7, 0.2])
now_ms = time.time() * 1000
print(round(budget, 1), freshness(now_ms - 5_000, budget, "open"),
freshness(now_ms - 60_000, budget, "open"), freshness(now_ms - 60_000, budget, "closed"))
# 3. Spike filter: a one-tick spike, then a real move that the next tick confirms
f = SpikeFilter()
for x in [100.00, 100.01, 99.99, 100.02, 100.00, 99.98, 100.01, 100.03, 100.00, 99.99,
100.02, 101.00, 100.01, 100.02, 102.50, 102.51, 102.49]:
result = f.check(x)
if result not in ("ok", "warming up"):
print(f"{x:.2f} {result}")None
crossed or locked
timestamp out of range (seconds instead of milliseconds?)
14.4 live stale closed
101.00 quarantined (z = 67.4)
102.50 quarantined (z = 167.9)
102.51 accepted: move confirmedThe window size, the z limit, the 0.1% confirmation band and the "three times the 99th percentile" budget rule are starting points for your own tuning, not recommendations and not anyone's production settings. The code runs on Python 3.9 or newer with the standard library only.
Layer 4: cross-checks between related data
- Trade against quoteA trade far outside the prevailing bid and ask is suspect. Allow a little slack, because fast markets print just outside the quote.
- REST against streamThe REST snapshot and the last streamed quote should agree. A persistent gap means one path is stuck.
- Related instrumentsAn index and an ETF that tracks it move together; a break in their usual ratio flags one of them.
- Against the previous closeA daily change far beyond anything in the symbol's history is more often a unit error than news.
Related instruments make the strongest check because they fail independently. On 28 September 2026 the US500 index quote was 7,697.93 to 7,701.77 and the US500ETF quote 767.13 to 767.19, a ratio near 10.04 that barely changes from one minute to the next. If it suddenly read 1.004 or 100.4, one of the two has a scale error, and you know which one by checking each against its own history. Choosing sources with this in mind is part of comparing financial data providers.
What to do with a bad tick once you catch it
Do
- Quarantine the tick and log the reason next to it.
- Keep raw data; apply filters when you read.
- Show the last good value with its timestamp and a stale label.
- Alert on the rate of rejections per symbol, not on single ticks.
Avoid
- Deleting raw ticks, which makes the filter impossible to audit.
- Forward-filling a frozen price as if it were live.
- Interpolating trades that never happened.
- Paging someone for every rejected tick.
Anything automated should refuse to act rather than guess. The guard in the AI trading bot tutorial does exactly that: a stale, crossed or wide quote produces a rejection with a reason, and the model is told why.
Questions
What is a bad tick in trading?
A bad tick is a price or quote that does not reflect the tradable market, such as an erroneous spike, a zero or crossed quote, a unit error or a frozen value. Charts and signals built on it show moves that never happened.
What is stale data in market data?
Stale data is a value that has stopped updating while the market is still trading. It looks valid, so the only way to catch it is to compare its timestamp with a freshness budget for that symbol and session.
How do you detect outliers in a price time series?
Measure each new value against a rolling median using the median absolute deviation, flag values with a large robust z-score, and confirm with the next tick before accepting a jump.
Should I delete bad ticks?
No. Keep the raw data, mark suspect ticks with a reason and filter on read, so you can audit and re-tune the filter later.
How old is too old for a price?
It depends on the instrument and the session. Learn each symbol's normal gaps during open hours, set a budget from a high percentile of them, and treat silence outside the session as closed rather than stale.