Explainer

Bad ticks and stale prices: how to detect them in market data

Most bad data is not dramatic. It is a price that stopped moving an hour ago and still looks perfectly reasonable on a chart.

On this page
  1. A field guide to bad ticks
  2. Layer 1: structural checks on every message
  3. Layer 2: stale data and freshness budgets per asset and session
  4. Layer 3: spike filters and outlier detection in a time series
  5. Layer 4: cross-checks between related data
  6. What to do with a bad tick once you catch it
  7. Questions

Key takeaways

  • A bad tick is a price or quote that does not reflect the tradable market: a spike, a zero, a crossed quote, a unit or decimal error, or a value that stopped updating.
  • Stale data is the hardest case because every field looks valid; only the timestamp, judged against a session-aware freshness budget, gives it away.
  • Silence is normal for a stock after the close and for a bond yield between daily observations, so freshness budgets must depend on asset class and session.
  • For spikes, use a rolling median and median absolute deviation rather than a mean and standard deviation, and quarantine suspect ticks instead of deleting them.
  • Cross-checks between trade and quote, REST and stream, and related instruments catch errors that no single series can reveal.

A bad tick is a price or quote that does not represent the tradable market: a print ten percent away from its neighbours, a zero bid, a bid above the ask, a price in cents where you expected dollars, or a stale value that quietly stopped updating. Catch them in layers: structural checks, session-aware freshness budgets, a spike filter on robust statistics, and cross-checks against related data.

No feed is immune, which is why these checks belong in your code as well as upstream. This explainer describes the approach and shows tested Python. It deliberately publishes no one's internal thresholds: the right numbers depend on your instruments and your tolerance, and the code learns them from your own history. For how we measure our own feed, see data quality.

A field guide to bad ticks

TypeWhat it looks likeTypical causeCaught by
SpikeOne print far from its neighbours, then backErroneous trade, parse error, odd printSpike filter
Zero or negativebid: 0 or price: 0Empty book side, a missing value defaulted to zeroStructural check
Crossed or lockedBid at or above the askThe two sides updated at slightly different momentsStructural check
Unit or scale errorA price 100 times too large or smallCents against dollars, a shifted decimalUnit table, ratio to previous close
Stale or frozenSame value, old timestampDisconnected source, stuck cache, closed marketFreshness budget
Out of orderA timestamp earlier than the last oneLate delivery, a replay after reconnectMonotonic check per symbol
Wrong time unitDates in 1970 or far in the futureSeconds read as milliseconds, or the reversePlausible-range check
DuplicateThe same trade twiceReplay after a reconnectDedupe on symbol, time, price and size

Unit errors deserve a concrete case. SUGARUSD is quoted in US cents per pound, so a snapshot bid of 18.537 means 18.5 cents, not $18.54. A pipeline that assumes dollars will either flag every sugar tick as a 99% crash or, worse, store it. Keep a unit table per instrument; the sugar price page shows the documented unit. Time units fail the same way: a Unix timestamp in seconds read as milliseconds lands in January 1970, and the Unix timestamp guide shows how to tell them apart.

Layer 1: structural checks on every message

Start with checks that need no history: every numeric field parses (WebSocket frames for crypto, forex and stocks send numbers as strings), both sides are positive, the ask is above the bid, the timestamp is in a plausible millisecond range, and it is not earlier than the last one you accepted for that symbol. These catch zeros, crossed quotes, unit mix-ups in time and out-of-order replays for almost no CPU. The bid-ask spread explainer covers why crossed and locked quotes happen.

Layer 2: stale data and freshness budgets per asset and session

Stale data is the dangerous case because nothing about the value looks wrong. The only signal is time: how long since this symbol last updated, compared with how long it normally stays quiet at this hour. That comparison needs two inputs, a budget per symbol and session, and the market status, because outside the session silence is expected and the right output is "closed, last value at 15:30", not an alarm.

AssetWhen silence is normalHow to set the budget
CryptoNever fully; nights and weekends are thinnerFrom your own quiet-hour gaps, per pair
ForexThe weekend close; thin minutes around the daily rolloverShort on weekdays, suspended at the weekend
US stocksOutside 04:00 to 20:00 New York, holidays, haltsPer symbol and per session; extended hours need a wider budget
CommoditiesThe daily 17:00 to 18:00 New York pause and the weekendSession-aware, like forex
IndicesOutside the underlying market's sessionFollow that market's calendar
Bond yieldsBetween daily observationsIn days: a Friday observation read on Monday is current
1-minute barsA minute with no trades produces no barNever alarm on a missing bar for a thin symbol
Session boundaries come from market status and calendar endpoints; see the [market status overview](/market-status).

When silence is expected, weekday in UTC

CryptoTrading
US stocksdaylight timePreRegularPost
Commoditiesmetals and energy

Hours in UTC

Gaps in a row are hours when a missing update is expected, not a bad tick. Extended sessions are thinner, so they need wider budgets.

Do not hard-code budgets from a blog post, including this one. Record the gaps between updates for each symbol during normal sessions, take a high percentile, multiply by a safety factor, and review the rejections. A liquid pair and a thin stock can differ by orders of magnitude, and so can the same stock in regular and extended hours.

Layer 3: spike filters and outlier detection in a time series

The textbook z-score, distance from the mean in standard deviations, fails exactly when you need it: a large spike inflates the standard deviation it is measured against. Use the median and the median absolute deviation (MAD) instead. They barely move when one value is wild, so the spike stands out.

robust z = |x − median(window)| ÷ (1.4826 × MAD(window))

window
The last N accepted values for this symbol, for example mids.
MAD
Median of absolute deviations from the median.
1.4826
Scales MAD to match a standard deviation on normally distributed data.
Illustrative: median 100.00 and MAD 0.02 put a print at 101.00 at a robust z of 1.00 ÷ 0.0297 ≈ 34, far outside any sane band.

A one-tick spike against the rolling median

  • Price
  • Rolling median
Illustrative series, not market data. The same filter drops the first jump and accepts the second because the next tick agreed with it.

The second trick is confirmation. A bad tick reverts; a real move persists. So hold a suspect value for one tick: if the next tick lands near the suspect, accept both and rebuild the window at the new level; if it lands back near the median, drop the suspect. You pay one tick of delay on genuine gaps, which is the right trade for anything that alerts or acts.

tick_checks.pyPython
from __future__ import annotations

import statistics
import time
from collections import deque

LAST_TS: dict[tuple[str, str], int] = {}


def structural_problem(msg: dict, kind: str = "quote") -> str | None:
    """Why a quote or trade frame is unusable, or None if it passes."""
    try:
        ts = int(msg.get("ts") or msg["timestamp"])
        if kind == "quote":
            bid, ask = float(msg["bid"]), float(msg["ask"])  # WebSocket numerics can be strings
        else:
            price = float(msg["price"])
    except (KeyError, TypeError, ValueError):
        return "missing or non-numeric field"
    if not 946_684_800_000 <= ts <= time.time() * 1000 + 5_000:
        return "timestamp out of range (seconds instead of milliseconds?)"
    key = (msg["symbol"], kind)
    if ts < LAST_TS.get(key, 0):
        return "out of order"
    if kind == "quote" and (bid <= 0 or ask <= 0):
        return "zero or negative side"
    if kind == "quote" and ask <= bid:
        return "crossed or locked"
    if kind == "trade" and price <= 0:
        return "zero or negative price"
    LAST_TS[key] = ts
    return None


def learn_budget(gaps_s: list[float], floor_s: float = 1.0) -> float:
    """Freshness budget from your own history: 3x the 99th percentile gap."""
    p99 = statistics.quantiles(gaps_s, n=100)[98]
    return max(floor_s, 3 * p99)


def freshness(ts_ms: int, budget_s: float, status: str) -> str:
    """Classify the last value of one symbol as live, stale or closed."""
    if status == "closed":
        return "closed"  # expected silence: show the last value and its time
    age_s = time.time() - ts_ms / 1000
    return "stale" if age_s > budget_s else "live"


class SpikeFilter:
    """Rolling median and MAD with one-tick confirmation."""

    def __init__(self, window: int = 50, z_max: float = 8.0, min_rel_scale: float = 1e-5):
        self.values: deque[float] = deque(maxlen=window)
        self.z_max = z_max
        self.min_rel_scale = min_rel_scale
        self.pending: float | None = None

    def _z(self, x: float) -> float:
        med = statistics.median(self.values)
        mad = statistics.median(abs(v - med) for v in self.values)
        scale = max(1.4826 * mad, abs(med) * self.min_rel_scale)
        return abs(x - med) / scale

    def check(self, x: float) -> str:
        if len(self.values) < 10:
            self.values.append(x)
            return "warming up"
        if self.pending is not None:
            suspect, self.pending = self.pending, None
            if abs(x - suspect) <= abs(suspect) * 0.001:  # next tick agrees with the jump
                self.values.clear()  # new price level: rebuild the window from here
                self.values.extend([suspect, x])
                return "accepted: move confirmed"
        if self._z(x) > self.z_max:
            self.pending = x
            return f"quarantined (z = {self._z(x):.1f})"
        self.values.append(x)
        return "ok"
tick_demo.pyPython
import time
from tick_checks import SpikeFilter, freshness, learn_budget, structural_problem

# 1. Structural checks on stream-style frames (numerics arrive as strings)
print(structural_problem({"symbol": "EURUSD", "bid": "1.137107", "ask": "1.137157", "ts": 1790591203002}))
print(structural_problem({"symbol": "EURUSD", "bid": "1.137157", "ask": "1.137107", "ts": 1790591204002}))
print(structural_problem({"symbol": "US:KO", "price": "88.24", "ts": 1790589160}, kind="trade"))

# 2. Freshness: budget learned from observed gaps (seconds), then applied
budget = learn_budget([0.3, 0.5, 1.2, 0.8, 2.0, 0.4, 0.6, 0.9, 1.1, 3.5, 0.7, 0.2])
now_ms = time.time() * 1000
print(round(budget, 1), freshness(now_ms - 5_000, budget, "open"),
      freshness(now_ms - 60_000, budget, "open"), freshness(now_ms - 60_000, budget, "closed"))

# 3. Spike filter: a one-tick spike, then a real move that the next tick confirms
f = SpikeFilter()
for x in [100.00, 100.01, 99.99, 100.02, 100.00, 99.98, 100.01, 100.03, 100.00, 99.99,
          100.02, 101.00, 100.01, 100.02, 102.50, 102.51, 102.49]:
    result = f.check(x)
    if result not in ("ok", "warming up"):
        print(f"{x:.2f} {result}")
Output
None
crossed or locked
timestamp out of range (seconds instead of milliseconds?)
14.4 live stale closed
101.00 quarantined (z = 67.4)
102.50 quarantined (z = 167.9)
102.51 accepted: move confirmed

The window size, the z limit, the 0.1% confirmation band and the "three times the 99th percentile" budget rule are starting points for your own tuning, not recommendations and not anyone's production settings. The code runs on Python 3.9 or newer with the standard library only.

  • Trade against quoteA trade far outside the prevailing bid and ask is suspect. Allow a little slack, because fast markets print just outside the quote.
  • REST against streamThe REST snapshot and the last streamed quote should agree. A persistent gap means one path is stuck.
  • Related instrumentsAn index and an ETF that tracks it move together; a break in their usual ratio flags one of them.
  • Against the previous closeA daily change far beyond anything in the symbol's history is more often a unit error than news.

Related instruments make the strongest check because they fail independently. On 28 September 2026 the US500 index quote was 7,697.93 to 7,701.77 and the US500ETF quote 767.13 to 767.19, a ratio near 10.04 that barely changes from one minute to the next. If it suddenly read 1.004 or 100.4, one of the two has a scale error, and you know which one by checking each against its own history. Choosing sources with this in mind is part of comparing financial data providers.

What to do with a bad tick once you catch it

Do

  • Quarantine the tick and log the reason next to it.
  • Keep raw data; apply filters when you read.
  • Show the last good value with its timestamp and a stale label.
  • Alert on the rate of rejections per symbol, not on single ticks.

Avoid

  • Deleting raw ticks, which makes the filter impossible to audit.
  • Forward-filling a frozen price as if it were live.
  • Interpolating trades that never happened.
  • Paging someone for every rejected tick.

Anything automated should refuse to act rather than guess. The guard in the AI trading bot tutorial does exactly that: a stale, crossed or wide quote produces a rejection with a reason, and the model is told why.

Questions

What is a bad tick in trading?

A bad tick is a price or quote that does not reflect the tradable market, such as an erroneous spike, a zero or crossed quote, a unit error or a frozen value. Charts and signals built on it show moves that never happened.

What is stale data in market data?

Stale data is a value that has stopped updating while the market is still trading. It looks valid, so the only way to catch it is to compare its timestamp with a freshness budget for that symbol and session.

How do you detect outliers in a price time series?

Measure each new value against a rolling median using the median absolute deviation, flag values with a large robust z-score, and confirm with the next tick before accepting a jump.

Should I delete bad ticks?

No. Keep the raw data, mark suspect ticks with a reason and filter on read, so you can audit and re-tune the filter later.

How old is too old for a price?

It depends on the instrument and the session. Learn each symbol's normal gaps during open hours, set a budget from a high percentile of them, and treat silence outside the session as closed rather than stale.

Keep reading

Ready to integrate?

Start with the free tier, explore the docs, and connect via REST or WebSocket in minutes.