Assay report

Assay

An assay tests a lump of metal to find out what it really contains. This page does the same to trading strategies. Every number here was measured, and the ones that can’t be trusted yet say so.

One year says it works.
The two before say it doesn’t.

Three versions of the Isaac model, which differ only in the hours of the day they may trade. Each bar is the average result of one trade: right of the centre line is a gain, left is a loss. Results are in R, where 1R is the amount a trade stood to lose, so −0.1R means the average trade lost a tenth of what it risked. The dashed outline shows the same rules over the other stretch of time.

What am I looking at?

Under each name is the number of trades and a chance test, which asks whether the result could easily be luck. The closer it is to zero, the more likely luck explains it.

These versions were picked by looking at the last twelve months, so that year cannot also be the proof. Try several versions on the same data and one will always come out on top. The two years before are the fair check, because nothing was picked using them. A real result would point the same way in both. Switch between them and see.

The fair test

Nine well-known sets of trading rules, each tested on one price per day over eighteen years. Most were written for daily prices, so this is the test their authors had in mind. Point at or tap an underlined column heading to see what it means.

What am I looking at?

Look at the chance test and the enough-data columns first. They tell you how far to trust the rest of the row. A chance test near zero means luck could explain the result. A short enough-data bar means there are not yet enough trades to be sure either way.

A high share of winning trades can still lose money overall, if the wins are small and the losses are big. And with this many rows, one is likely to look good by luck alone, so a single promising number here is worth less than it seems.

Does adding to winning trades help?

Tom’s rules don’t buy everything at once. When a trade starts going the right way they add more, up to four times, and one exit closes the lot. Traders call this pyramiding. It is where this style of trading is supposed to make its money, so it had to be tested before anyone could fairly say Tom’s rules lose.

Most held at onceTrades Average, to scale

Adding to a trade makes it bigger, not better. If the rules have an edge, adding makes the gains bigger. If they don’t, it makes the losses bigger, because you keep buying more of a move that then fades.

What am I looking at?

Each row is the same rules with a different limit on how many additions are allowed.

The number of trades changes from row to row. Adding more changes when a trade closes, and so when the next one can start. The rows are related, but they are genuinely different sets of trades.

All eleven, side by side

Every rule set gets the same currency pairs, the same dates and the same trading costs, using a price every 15 minutes. Drag the slider to change the trading cost and watch the results move. A result that only looks good when trading is almost free isn’t a real result.

This is a fair comparison between them, not the proper test of each. Nine of the eleven were written for daily prices, and their proper test is the fair test above. Only Isaac and Sam were written for fast prices like these.

0.8 pips

What am I looking at?

The slider sets the trading cost, called the spread. It is the small gap between the price you can buy at and the price you can sell at, and every trade pays it, win or lose.

Click a row to see how its results built up over time. An average hides the shape: a slow, steady loss and one terrible month can give the same average, and they feel very different to live through.

What the agent is deciding

This part is about the software, not the strategy. It shows how many trade ideas the agent looked at, how many it turned down, and what that cost. Turning most of them down is correct: the plan describes a rare situation.

What am I looking at?

The breakdown separates an ordinary refusal from one where a must-pass check failed.

How long alerts take to answer is shown as a range of delays, never as times of day. Times of day would show when the owner is awake, which is nobody else’s business.

Which checks turn ideas down

Before an idea becomes an alert, the AI goes through the plan’s checklist and scores each item from 0 to 100. The final score is worked out by the software, not the AI. An item that always scores 0 suggests the software and the AI disagree about what it means. One that often lands in the middle suggests the AI is unsure about a yes-or-no question.

Scored 0 1 to 49 50 to 79 80 or more
Checklist itemTimes scoredAverage scoreTimes it blockedHow it scored
What am I looking at?

For every score the AI also says what decided it. Because the software works out the total, the AI cannot talk a trade up.

Some questions are marked must pass. If one of those scores too low, the idea is turned down whatever the other answers say. The blocked column counts how often that happened.

The pattern of scores matters most. A simple count of blocks cannot tell the two problems above apart, and they need different fixes.

This table is about the judge, not the strategy.

What the words mean

Every term used on this page, in plain English. The underlined column headings above show the same explanations when you point at them or tap them.

Words on this page
R
The amount a trade stood to lose. Every trade is planned with an exit point that admits it was wrong, and the distance to that exit is what the trade risks. Measuring results in R lets big and small trades be averaged fairly.
Average per trade
The average result of one trade, in R. Above zero, the rules made money on average; below zero, they lost. Traders call this expectancy.
Total
All the trades added together, in R.
Winners
The share of trades that made money. It says nothing about how much, so a high share can still add up to a loss.
Profit factor
Everything won divided by everything lost. Above one means ahead overall; below one means behind.
Chance test
Whether a result could easily be luck. It compares how far the average is from zero with how much the individual trades varied. Near zero, luck explains it easily. Because so many rule sets are compared at once, this page wants something close to three, either side of zero, before it takes a result seriously. Statisticians call this the t statistic.
Break-even cost
The trading cost at which a rule set stops making money. Never means it loses even when trading is free, so no cheaper broker could rescue it.
Enough data?
The trades on hand compared with the number a result of that size would need before it could be trusted. A bar a fifth full means the honest answer is still not yet.
Verdict
A one-line summary of the row, judged against the stricter bar the chance test describes. Loses: it loses clearly, even before trading costs. Fee eats it: it loses clearly, but only because of trading costs. No evidence: too close to zero to tell from luck. Clears the bar: a clear gain, even after costs. A verdict never claims more than the row beside it.
Spread
The small gap between the price you can buy at and the price you can sell at. It is the main cost of trading, and every trade pays it, win or lose. It is measured in pips.
Pip
The smallest usual step in a currency’s price. For most pairs it is one ten-thousandth of the price.
Currency pair
Two currencies priced against each other, such as the Japanese yen against the Swiss franc.
Kill zone
A few busy hours around the opening of the London or New York markets. Some of these rules only trade inside one.
Daily, fifteen-minute
How much time each price covers: one price per day, or one every fifteen minutes. Rules written for one speed behave differently at another.
Testing on fresh data
Choosing rules on one stretch of history, then checking them on a different stretch. Anything that only works on the stretch it was chosen from has not really been tested.