← Back to the site

How we test a trading idea

What happens between "describe your idea" and the verdict — and what we do not do yet. Every line on this page describes the engine as it runs today.

1. Costs: we charge what a real account pays

Commission of $1 per side on a $1,000 reference position (0.2% round trip).

The bid-ask spread, by the minute. Half the spread is charged on entry and half on exit, each at the minute the trade happens. The spread depends on the share price, the stock's dollar volume and the time of day: it is widest in the first 15 minutes after the open and about five times narrower near the close. The rates come from 1,879 real quotes, and we charge the conservative end (75th percentile), not the average. An exchange access fee is added per share.

Short borrow, per day held. Easy-to-borrow stocks pay 0.03% of the position per day, hard-to-borrow 0.50%, very hard 1.70%. A short on a stock that already moved 40% or more is charged at least 3%.

Most ideas that look profitable before these costs stop being profitable after them. That is the point of charging them.

2. A period the search does not train on

The best settings are searched on one period (by default 2020–2024) and then checked on a later period (by default 2025 to today). An idea "holds" only if the later period keeps at least half of the earlier result, on at least 20 trades.

Being exact about it: in our gap engine, a few later steps (accepting an enhancement, choosing between close variants) do look at the held-out result. In the multi-day engine the winner is chosen on the training part only. We would rather say this than claim a purity we do not have.

3. Stocks that no longer exist are included

Our data includes delisted stocks (about 34,600 symbols back to 2016). Testing only on stocks that still trade today makes almost every long strategy look better than it was. After each test an automatic check flags a trade list made only of survivors.

4. No peeking at the future

A rule may only use information known at the moment of the trade. A test that wants to enter at the open on a signal that is only known at the close is refused, with the reason.

After every run we also check that what ran is what you asked for (sector, universe, side, timing) and flag any gap, and we run a list of checks on the trades themselves: traded prices, borrow, calendar, overlapping positions, price floor, regular hours only.

5. Results that look too good are stopped

A win rate of 80%+, 8%+ per trade, or a held-out period far better than the training period is treated as a sign of a modeling gap on our side, not as a discovery. Such a result is re-run with four times the slippage; if the edge does not survive, you do not get the result, you get an explanation, and the test is not charged.

6. We refuse what we cannot measure

If an idea needs data we do not carry (for example FDA decisions, clinical trial results, M&A, analyst calls, or option prices), we say so and do not run it. Running it on other data would give a confident, wrong answer.

Penny stocks are excluded (minimum price $1–$5 depending on the engine), and so are warrants, units and rights.

7. Overfitting checks

Plateau, not spike: the chosen settings must sit among neighbours that also work. Year by year: stable, mostly stable, unstable or decaying. Recent months: checked separately. Search size: if the best result is about what trying that many combinations would produce by chance, a "promising" rating is downgraded.

The rating is promising, needs work or weak; any one of these warnings keeps an idea out of "promising".

8. After the test: tracked every day

Every saved strategy is re-run each weekday after the close on the new days only, starting from the day it was saved — so no backtest day is counted twice. These are simulated trades on new data, not real orders. A month-by-month copy of this record is kept, so it cannot be quietly rewritten later.

What we do not do yet

No walk-forward analysis, no permutation tests and no formal multiple-testing correction in the main engine. The deflated Sharpe ratio is shown for information but does not change a verdict. Some single-symbol and IPO tests use a flat cost instead of the per-minute spread.

A backtest, however careful, is not a promise. It tells you whether an idea is worth watching, not whether it will make money.

The data

Daily and hourly bars from 2016, one-minute bars from 2020, news headlines from 2021, and company earnings report dates from SEC filings. Market data comes from Alpaca.

See it on your own idea

Describe it in a sentence. No account needed.

Test an idea