Last updated: 2026-08-08
ml4t-backtest is designed to express external backtesting behavior through configurable profiles and to measure every remaining difference.
There are NO "expected differences." Every trade count gap, every value gap, is a signal that a configurable knob is missing or misconfigured. The gap must be driven to zero.
Every behavioral choice that differs between backtesting frameworks is expressed as a configurable
parameter in BacktestConfig. These are not workarounds or compatibility shims -- they are
first-class, well-documented behavioral dimensions that represent real design choices.
Examples of behavioral dimensions:
- Execution timing: same-bar vs next-bar
- Fill price: open, close, VWAP, midpoint
- Stop level basis: fill price vs signal price
- Cash policy: constrained vs unconstrained, credit vs lock_notional
- Fill ordering: exit-first vs FIFO
- Share type: integer vs fractional
- Commission model: none, percentage, per-share
- Rebalance mode: incremental vs snapshot
Each external backtester gets a profile that sets ALL configurable knobs to the values that replicate that framework's behavior:
| Profile | Emulates | Goal |
|---|---|---|
vectorbt / vectorbt_strict |
VectorBT Pro/OSS | 0% trade gap, 0% value gap |
backtrader / backtrader_strict |
Backtrader | 0% trade gap, 0% value gap |
zipline / zipline_strict |
Zipline Reloaded | 0% trade gap, 0% value gap |
lean |
QuantConnect LEAN | 0% fill gap, near-0% value gap |
default |
ml4t's own opinion | Best-practice defaults |
realistic |
Conservative simulation | Adds costs, integer shares |
The default profile represents ml4t's own opinion on the most reasonable settings:
| Dimension | ml4t Default | Rationale |
|---|---|---|
| Execution | next_bar_open | No look-ahead bias |
| Costs | 0.1% commission + 0.1% slippage | Conservative but not punishing |
| Short selling | disabled | Require explicit opt-in |
| Leverage | disabled | Require explicit opt-in |
| Share type | fractional | Simpler for research |
| Fill ordering | exit_first | Capital-efficient |
| Rebalance | incremental | Most accurate cash tracking |
| Stop basis | fill_price | Based on actual execution |
Every choice is documented and justified. Users can see exactly what the default profile does and why, and they can override any setting.
The configuration system makes ALL behavioral choices explicit and visible:
- A user can inspect any profile and see exactly which settings differ from default
- The behavioral difference matrix (below) documents what each framework does
- No hidden behaviors -- if a framework does something differently, it's a named config parameter
Convergence with independently developed frameworks provides evidence for the execution behavior covered by the retained workloads. It does not establish correctness for behavior that the comparison matrix does not exercise.
Performance comparisons are valid only after the profile reproduces the behavior exercised by the benchmark and every framework runs the same retained workload.
Users migrating from another framework can inspect retained comparison results for the matching profile before switching to ml4t's default or realistic profiles.
Run the benchmark suite against an external framework. Any gap -- no matter how small -- indicates a behavioral difference that needs to be captured.
For each gap, identify the specific behavioral dimension that differs. This is NOT "debug the bug" -- it's "which design choice does this framework make differently?"
Common root causes:
- Fill price timing (close vs open vs next-bar open)
- Cash constraint model (constrained vs unconstrained)
- Order processing sequence (exit-first vs FIFO)
- Position sizing (integer vs fractional shares)
- Commission/slippage application
If the behavioral dimension isn't already configurable in BacktestConfig, add it:
- Add the parameter to
BacktestConfigwith a sensible default - Wire it through the execution path (broker, engine, accounting)
- Add tests for both values of the parameter
- Document it in the behavioral difference matrix
Set the new parameter in the framework's profile to match its behavior.
Run the benchmark suite again. The gap for this dimension should now be zero. If not, there are additional behavioral differences to capture -- go back to step 1.
Release comparisons use equality after all adapters serialize numeric values to an eight-decimal fixed-point representation. This representation removes framework-specific binary floating-point noise without allowing a nonzero value difference. Comparison artifacts retain raw values, raw differences, canonical values, record counts, hashes, and the first divergent record. Scenario diagnostic thresholds provide context in failure messages and never change pass or fail status.
Update the behavioral difference matrix, profile diffs, and parity results.
DO NOT accept gaps as "structural" or "expected":
- "LEAN fills at bar close, ml4t at next-bar open -- this is a structural difference" is WRONG
- CORRECT: "We need to add
execution_price=closeto the lean profile to close this gap"
DO NOT say "this is good enough for production":
- 95% parity is not the goal -- 100% parity is the goal
- Every trade difference represents a behavioral dimension we haven't captured yet
DO NOT compare frameworks with mismatched profiles:
- Comparing ml4t[default] vs Backtrader is meaningless
- Always compare ml4t[backtrader_strict] vs Backtrader
Complete matrix of how each framework handles every behavioral dimension.
| Knob | ml4t Default | VectorBT | Backtrader | Zipline | LEAN |
|---|---|---|---|---|---|
execution_mode |
next_bar | same_bar | next_bar | next_bar | same_bar |
fill_timing |
next_bar_open | same_bar | next_bar_open | next_bar_open | same_bar |
execution_price |
open | close | open | open | close |
VectorBT and LEAN both use same-bar execution with close fills. VectorBT is vectorized (inherently same-bar), while LEAN processes market orders immediately at bar close in its event-driven loop. All event-driven frameworks (BT, Zipline) except LEAN use next-bar open.
| Knob | ml4t Default | VectorBT | Backtrader | Zipline | LEAN |
|---|---|---|---|---|---|
stop_fill_mode |
stop_price | stop_price | stop_price | stop_price | stop_price |
stop_level_basis |
fill_price | fill_price | signal_price | fill_price | fill_price |
trail_hwm_source |
close | bar_extreme | close | close | close |
initial_hwm_source |
fill_price | bar_high | fill_price | fill_price | fill_price |
trail_stop_timing |
lagged | intrabar | lagged | lagged | lagged |
Backtrader calculates stop levels from signal bar close (the price when the strategy decided to trade), not the actual fill price. This matters when next-bar open differs significantly from previous close.
| Knob | ml4t Default | VectorBT | Backtrader | Zipline | LEAN |
|---|---|---|---|---|---|
allow_short_selling |
false | true | true | false | true |
allow_leverage |
false | false | true | false | true |
short_cash_policy |
credit | credit | credit | credit | credit |
initial_margin |
0.5 | -- | 0.5 | -- | varies |
long_maint_margin |
0.25 | -- | 0.25 | -- | varies |
short_maint_margin |
0.30 | -- | 0.30 | -- | varies |
| Knob | ml4t Default | VectorBT | Backtrader | Zipline | LEAN |
|---|---|---|---|---|---|
fill_ordering |
exit_first | exit_first | fifo | exit_first | exit_first |
entry_order_priority |
submission | submission | submission | submission | submission |
rebalance_mode |
incremental | hybrid | snapshot | snapshot | snapshot |
rebalance_headroom_pct |
1.0 | 1.0 | 0.998 | 0.998 | 0.998 |
reject_on_insufficient_cash |
true | false | true | true | true |
partial_fills_allowed |
false | true | false | true | false |
missing_price_policy |
skip | use_last | use_last | use_last | use_last |
| Knob | ml4t Default | VectorBT | Backtrader | Zipline | LEAN |
|---|---|---|---|---|---|
share_type |
fractional | fractional | integer | integer | integer |
signal_processing |
check_position | process_all | check_position | check_position | check_position |
commission_model |
percentage | none | percentage | per_share | per_share |
commission_rate |
0.1% | 0% | 0.1% | -- | -- |
commission_per_share |
-- | -- | -- | $0.005 | varies |
commission_minimum |
$0 | -- | -- | $1.00 | varies |
slippage_model |
percentage | none | percentage | volume_based | volume_based |
slippage_rate |
0.1% | 0% | 0.1% | 10% | varies |
| Knob | vbt_strict | bt_strict | zip_strict | lean |
|---|---|---|---|---|
short_cash_policy |
lock_notional | credit | credit | credit |
skip_cash_validation |
false | false | true | false |
reject_on_insufficient_cash |
true | true | true | true |
fill_ordering |
fifo | fifo | exit_first | exit_first |
next_bar_submission_precheck |
-- | true | -- | false |
next_bar_simple_cash_check |
-- | true | -- | false |
allow_short_selling |
-- | -- | true | true |
allow_leverage |
-- | -- | -- | true |
execution_mode |
-- | next_bar | next_bar | next_bar |
execution_price |
-- | open | open | open |
Scenario claims use the retained release-candidate matrix. "Exact" appears only when every required scenario has zero canonical gap.
| Profile | Pinned framework | Required scenarios | Evidence |
|---|---|---|---|
vectorbt_strict |
VectorBT Pro 2025.12.31 | 16/16 exact | scenario evidence |
vectorbt |
VectorBT OSS 0.28.2 | 15/15 exact | scenario evidence |
backtrader_strict |
Backtrader 1.9.78.123 | 16/16 exact | scenario evidence |
zipline_strict |
Zipline Reloaded 3.1.1 | 15/15 exact | scenario evidence |
Large-scale claims are published only when a retained workload has zero canonical gap.
| Profile | Pinned framework | Compared | Trade gap | Terminal value | Evidence |
|---|---|---|---|---|---|
vectorbt_strict |
VectorBT Pro 2025.12.31 (1305a1e19743) |
225,844 trades | 0 | 685179.007330 | large-scale evidence |
No large-scale claim is published for Backtrader, Zipline, VectorBT OSS, or LEAN without a passing retained artifact.
The validation suite tests 16 scenarios against 4 external frameworks:
| ID | Scenario | What It Tests |
|---|---|---|
| 01 | Long Only | Basic execution, fill price |
| 02 | Long/Short | Direction switching, short selling |
| 03 | Stop Loss | Risk rule triggering, stop fill mode |
| 04 | Take Profit | Limit order execution |
| 05 | Commission (Pct) | Percentage commission model |
| 06 | Commission (Per-Share) | Per-share commission model |
| 07 | Slippage (Fixed) | Fixed slippage model |
| 08 | Slippage (Pct) | Percentage slippage model |
| 09 | Trailing Stop | High-water-mark tracking, exit timing |
| 10 | Bracket Order | OCO orders, rule chain priority |
| 11 | Short Only | Short position mechanics |
| 12 | Short Trailing Stop | Short-side high-water-mark |
| 13 | TSL + TP Combo | Rule chain evaluation order |
| 14 | TSL + SL Combo | Rule chain evaluation order |
| 15 | Triple Rule | Three-rule chain with priority |
| 16 | Stress Test | 1500 bars, 9 market regimes |
Beyond scenario tests, the benchmark suite runs a full ranking strategy on 250 US equities
from 1998-2018. This produces 200K+ trades that are compared trade-by-trade between frameworks.
See benchmark_suite.py for the implementation.
# Single scenario
python validation/run_scenario.py --scenario 01 --framework backtrader
# All scenarios for one framework
python validation/run_scenario.py --framework vectorbt_oss
# Full matrix
python validation/run_scenario.py --all
# Large-scale benchmark
python validation/benchmark_suite.py --profile backtrader_strict --framework backtrader| Resource | Path |
|---|---|
| Config (40+ knobs) | src/ml4t/backtest/config.py |
| Profiles (6 core + 4 strict) | src/ml4t/backtest/profiles.py |
| Scenario definitions | validation/scenarios/definitions.py |
| Framework drivers | validation/frameworks/ |
| Benchmark suite | validation/benchmark_suite.py |
| Scenario runner | validation/run_scenario.py |