DeskTested

Trading tools, tested on a real desk.

Updated Aug 23, 2026Desk NotesClaims verified ✓Anthony M.

How to Test a Trading Tool Before You Recommend It

The trial protocol this desk runs before any tool gets written about: a falsifiable hypothesis, a fixed method, recorded metrics, and a public log.

To test a trading tool properly you write down a falsifiable hypothesis about it before you open the software, fix the exact procedure and the metrics you’ll record, run it for a pre-committed length of time, and log observations with dates as you go. If the tool’s output can be independently recomputed against data you already hold, the test is stronger still — because then the trial can fail, and a test that cannot fail is marketing.

This page publishes that protocol before we have any results to show. As of 23 August 2026, our trials directory holds the blank template and exactly one open trial: TradingView, started 2026-08-23 and pre-committed to run 30 calendar days to 2026-09-22. Its observation log currently has zero rows. So the honest position is narrower than “we test tools” and wider than “we have never started”: one trial is running, none has finished, and this desk has published no first-hand tool findings. We state it that precisely because a protocol is only worth anything if the empty state is as visible as the full one — including the difference between started and completed, which is exactly the distinction a tool review is most tempted to blur.

Why the protocol comes before the review

Almost every “we tested it” claim in the trading-tools space is unfalsifiable. It describes features, quotes a pricing page, and reports a favorable impression. There is no hypothesis, so there is no way the trial could have gone badly, and therefore no information in the fact that it went well.

The fix is not more rigor in the write-up. It’s a commitment made in advance:

  1. State what would make the tool fail. If you cannot describe an outcome that would stop you recommending it, you are not testing anything.
  2. Fix the method before you start. Settings, universe, and what gets recorded — all decided up front, so results can’t be produced by quietly changing the procedure until the numbers cooperate.
  3. Pre-commit the duration. Otherwise the trial ends when the results look good.
  4. Log with dates. A dated observation trail is what separates a test from a recollection.

Everything below is the operational form of those four commitments.

The trial record: five fields plus a log

Every trial on this desk is opened as a file in pipeline/desk-log/trials/ from a fixed template. Five header fields and one running table:

Hypothesis. One sentence, phrased so it can be wrong. Not “evaluate the screener” but a specific claim about what the tool will do. The test of a good hypothesis is that you can describe today the observation that would refute it.

Method. The exact repeatable procedure: which preset or settings, which universe of instruments, and what gets recorded on each pass. Specific enough that someone else could execute it and get comparable output.

Start and duration. A start date and a fixed run length, both written down before the first observation.

Metrics. What is measured — candidates surfaced per day, false-positive rate, agreement with an independently computed benchmark. Chosen in advance, because metrics chosen afterward are chosen to flatter.

Log. A dated table of observations: date, observation, supporting data. Appended as the trial runs, never reconstructed at the end.

The template is deliberately short. A heavy protocol produces no completed trials, and a trial that doesn’t finish teaches nothing.

What makes a hypothesis testable

The strongest hypotheses compare the tool’s output to something you can compute yourself. That turns a subjective impression into an arithmetic check.

A worked example of the shape, drawn from a candidate hypothesis our week-34 log drafted for TrendSpider — a tool no trial has been opened on: “TrendSpider’s automated trendline and moving-average detection on daily SPY/QQQ charts reproduces the 200-day-SMA levels our own computation produces.”

That works as a hypothesis for a specific reason. This desk already computes 200-day averages on those two tickers from its own data pull, so the comparison has a right answer that exists independently of the tool. Either the levels reconcile or they don’t, and the trial can return an unwelcome result.

The trial actually open on this desk is the TradingView one, and its hypothesis is built the same way — reliability, screener output against our existing watchlist process, and alert-delivery latency, each recorded daily against pre-committed metrics rather than summarised at the end.

Nothing about either tool is asserted here. A hypothesis is a question, not a finding, and until a trial’s log is filled in and published this desk has nothing to report on either one.

The validation standards a credible trial has to meet

Where a trial involves a tool’s backtesting or signal-generation output, four standards apply. They are not specific to any product; they’re what separates a result from a curve-fit.

Test out of sample

Any rule tuned on a stretch of history will describe that stretch well. The only meaningful question is how it behaves on data it was not fitted to. Hold back a period, or step forward in time re-fitting as you go, and report that — not the in-sample figures the tool shows by default.

Insist on a real sample size

A handful of trades tells you nothing about a rule, because the variance swamps the signal. Treat roughly 100 trades as a floor for taking a result seriously at all, and understand that a floor is not a comfortable sample — it is the point below which the exercise is not worth doing. Fewer trades over a longer period is not a substitute; the count is what constrains the uncertainty.

Model the costs that actually apply

A backtest with no commissions, no spread, no slippage, and fills at the exact close is not a simulation of trading. It is a simulation of a market that doesn’t exist. Whatever the tool assumes by default, find the setting, change it to something realistic for your instrument and size, and re-run. Strategies that survive the first run and die on the second are common, and that outcome is the test working.

Forward test before capital

The last step is running the rule on live data, in real time, without money on it, for a pre-committed period. It is the only stage that catches the failures a historical test structurally cannot — data that arrives late or revised, signals that fire at times you can’t act on, and the gap between what you said you’d do and what you do when it’s happening.

What we will and won’t say about a tool

Three rules govern how a trial becomes a page.

No first-hand claim without a log. “We ran it” is only publishable when a dated trial file supports it. Where no completed trial exists, we commit to writing about the tool third-person from the vendor’s own published material and to saying on the page that no finished trial stands behind it. That is the standard this page sets going forward, and the trial logs are what make it checkable rather than merely stated.

Vendor specifics get verified at the time of writing. Prices, fees, and feature availability change without notice, so any such figure is checked against the vendor’s live page in the session the page is written — or it is described as a range with no number attached.

A negative result gets published. A protocol that only surfaces trials confirming the recommendation is a filter, not a method. The log is the deliverable, and it’s public whichever way it comes out.

Frequently asked questions

How long should a trading tool trial run?

Long enough to cover varied conditions, and fixed before you start. The number matters less than the pre-commitment — a duration chosen while you can already see the results is not a duration, it’s an exit.

What sample size does a backtest need to be credible?

Treat about 100 trades as the minimum for a result to carry any weight, and prefer considerably more. Below that, normal variance produces impressive-looking equity curves from rules with no edge at all.

Can you test a trading tool on a free trial?

Often yes — a free tier is enough to run a reconciliation-style hypothesis, where you’re checking whether the tool’s output matches something you compute independently. It’s usually not enough for a forward test, which needs the pre-committed duration to run uninterrupted.

What’s the difference between a tool trial and a strategy backtest?

A trial evaluates the instrument: does the software do what it claims, accurately and repeatably. A backtest evaluates a rule. Conflating them is how a good tool gets blamed for a bad strategy, and how a bad tool escapes scrutiny because the strategy happened to work.

Why publish a trial log with no results in it yet?

Because the alternative is publishing the protocol only once it has produced flattering results, which tells a reader nothing about how often it produces unflattering ones. One trial is open as of 23 August 2026 and it has recorded zero completed runs. Those counts will change, and the public record of them changing is the evidence — a log that only ever appears already full is not a log, it’s a brochure.


Educational content only. Nothing here is a recommendation to buy or sell any security or to purchase any product. No tool named above has a completed trial behind it, and no finding about any of them is reported here.