This document outlines the standards and best practices for contributing to the OdyssNet project — whether you are adding example scripts, improving the library, or fixing bugs.
OdyssNet relies on a highly modular library structure. To ensure long-term maintainability and performance, all contributions must adhere to these guidelines.
# Clone the repository
git clone <repo-url> && cd odyssnet
# Install in development mode (required — makes `from odyssnet import ...` work everywhere)
pip install -e ".[dev]"
# Verify installation
python -m pytest tests/
# OR simply
pytest tests/ # but not recommended as it will not see your env setupCUDA note:
requirements.txtpins CUDA 11.8. For RTX 4000/5000 series GPUs, install PyTorch separately:pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
odyssnet/ # Library source code
core/network.py # OdyssNet model
training/trainer.py # OdyssNetTrainer
training/chaos_optimizer.py # ChaosGrad zero-config optimizer
utils/ # Data, checkpointing, neurogenesis, history
tests/ # Test suite (mirrors odyssnet/ structure)
examples/ # Core validation scripts (identity, XOR, MNIST)
advanced/ # Complex experiments (reasoning, generation, transfer)
We distinguish between Core Validations and Feature Experiments.
- Purpose: Contains minimal, "hello world" style scripts that validate the core laws of OdyssNet physics.
- Examples:
convergence_identity.py(Can signals pass?),convergence_gates.py(Can it solve XOR?). - Rule: Scripts here should be extremely simple, fast, and prove a fundamental property of the architecture.
- Purpose: Contains complex tasks, task-specific logic, and demonstrations of advanced cognitive behaviors.
- Examples:
convergence_detective_thinking.py(Reasoning),convergence_latch.py(Willpower),convergence_hive_mind.py(Collective memory across bodies). - Rule: If you are building a task (like adding numbers, generating waves, or playing a game), it goes here.
⛔ DO NOT re-invent the wheel. ✅ DO use the Library.
Never write your own manual PyTorch training loop (optimizer.step(), loss.backward(), etc.) unless absolutely necessary for low-level research.
- Why? The
OdyssNetTrainerhandles:- Automatic Mixed Precision (AMP): Faster training on Tensor Cores.
- Gradient Accumulation: Simulating large batches.
- Ghost Gradients (Persistence): Advanced stabilization.
- State Management: Resetting hidden states automatically.
# ❌ BAD: Manual Loop
output = model(input)
loss = criterion(output, target)
loss.backward()
optimizer.step()
# ✅ GOOD: Trainer
trainer = OdyssNetTrainer(model, device='cuda')
trainer.train_batch(input, target, thinking_steps=10)If you need a new feature (e.g., a new loss function or a custom metric), extend OdyssNetTrainer or pass arguments to it. If the library is missing a critical feature, implement it in the library first, then use it in your example.
OdyssNet is sensitive to initialization. The default weight_init='resonant' is the recommended starting point for all tasks — it places the weight matrix at the Edge of Chaos (ρ(W) = 1.0) from the start and works across all network sizes.
For most tasks without specific constraints, use the default resonant initialization.
- Activation:
'tanh' - Weight Init:
'resonant'(Default) — Rademacher ±1 skeleton + spectral normalization to ρ = 1.0. Ensures signal fidelity without exploding or vanishing. Projection layers (embed/proj/decoder) automatically use'quiet'init. - Gate:
None(Default) — resolves to['none', 'none', 'identity'](memory identity gate enabled; starts closed with zero gate init and opens if training needs it). - Dropout:
0.0(Default) — Enable explicitly (e.g.,0.1) only when overfitting is observed.
model = OdyssNet(..., activation='tanh') # weight_init='resonant' is already the defaultIf resonant convergence is too slow on very small circuits:
- Activation:
'gelu'— Better gradient flow in sparse/small graphs. - Weight Init:
'xavier_uniform'— High variance ensures signals don't die in small circuits. - Gate (Optional):
'sigmoid'— Stronger branch control if needed. - Dropout:
0.0— Every neuron is vital in small networks.
model = OdyssNet(..., activation='gelu', weight_init='xavier_uniform', dropout_rate=0.0)
trainer = OdyssNetTrainer(model, ..., synaptic_noise=0.0) # Disable noise for pure logicFor long-horizon temporal stability:
- Activation:
'tanh' - Weight Init:
'orthogonal'— Solid fallback for pure stability. - Gate (Optional):
['none', 'none', 'sigmoid']— Memory-only gating. - Dropout:
0.0(Default) — Enable explicitly when overfitting is a concern.
model = OdyssNet(..., activation='tanh', weight_init='orthogonal')gate=None: Default branch layout['none', 'none', 'identity'].gate='sigmoid': Applies same gate activation to all[encoder_decoder, core, memory]branches.gate=['none', 'none', 'sigmoid']: Only memory branch is gated.gate=['none', 'none', 'none']: Disables all gating.- List supports 1-3 entries, right-padded from defaults.
'none': Disables gate branch entirely (no learnable parameters).'identity': Enables explicit identity gating (learnable gate params exist, starts at identity).- Gate parameter initialization uses the 4th
weight_initslot. Default layout is['quiet', 'resonant', 'quiet', 'zero', 'quiet']— the 5th slot is the attention projections (3.0), and shorter lists are right-padded from the defaults, so a four-entry list written before 3.0 still means what it meant. - Activation layout supports 1-4 entries with default
['none', 'tanh', 'tanh', 'none']; 4th slot is reserved for config symmetry.
For tasks requiring precise storage and retrieval of values over time (e.g., Neural Database):
- Structure: High neuron count (256+) to provide storage space for memories.
- Init: Default
resonantinitialization is appropriate.
model = OdyssNet(num_neurons=256, ...) # resonant default works wellFor tasks requiring high input/output dimensionality (like vision or LLMs) without scaling the core state size:
- Feature: Use
vocab_size=(V_IN, V_OUT)to decouple input/output resolution from internal neuron count. - Optimization: Allows a tiny "Thinking Core" (e.g., 10 neurons) to process high-resolution signals (e.g., 784 pixels), achieving extreme parametric efficiency.
- Usage: Best used in conjunction with sequential signal processing for maximum compression.
- Note: When
weight_init='resonant', projection layers (embed/proj/decoder) automatically use'quiet'init (Normal(0, 0.02)) — no manual override needed.
# OdyssNet core has N=10 neurons, but processes 784 input channels and 10 output classes
model = OdyssNet(num_neurons=10, ..., vocab_size=(784, 10))For tasks where online synaptic plasticity may help — e.g., fast-adaptation, continual learning, or tasks with shifting statistics — enable hebb_type ("temporal", "spatial", or "both") and choose one of the three resolution levels:
hebb_res |
Extra Params per path | When to use |
|---|---|---|
"global" |
+2 | Quick experiments; uniform plasticity across all synapses. |
"neuron" |
+2N | RL and reactive environments; per-neuron "caste" differentiation. |
"synapse" |
+2N² | Logic, NLP, and reasoning tasks requiring dynamic variable binding. |
-
What it does: At each step the network accumulates correlations (temporal
$C_t = h_t \otimes h_{t-1}$ and/or spatial$C_s = h_t \otimes h_t$ ) depending onhebb_type. It applies them to the weights. Both factors and decays (t_hebb_factor,s_hebb_factor, etc.) are learnable — the network discovers how plastic each synapse should be. -
State: Correlations are persisted via buffers (
t_hebb_state_W,s_hebb_state_mem, etc.) across intra-sequence forward calls and are explicitly cleared onreset_state()between sequences. - Best Use Case (Generation / Sequential Building): Hebbian shines in tasks where step T relies heavily on expanding or completing a pattern from step T-1. It provides a powerful short-term working memory between steps, acting as a dynamic shortcut that fast-tracks sequence generation.
- When not to use it (Classification / Independent Features): Avoid Hebbian in classification tasks where each step processes distinct, independent chunks of information (e.g. sequential MNIST classification). In these tasks, inter-step short-term memory acts as "overfit noise".
-
Compatibility: Fully compatible with
gradient_checkpointing=True. - Combined with gating: Hebbian and gate parameters are independent groups; both can be active simultaneously.
-
Free to try: the plastic contribution is scaled by a zero-initialized gain (
hebb_norm, +N parameters) and the module draws no RNG, so a plastic model and a plain one at the same seed share a core and agree exactly until training moves the gain. Ablations ofhebb_typehave one variable in them. -
Cost: the trace is per example and is kept as the writes it is made of rather than as a matrix, so memory grows as
steps x B x Nand no(B, N, N)tensor is ever assembled. It is still the one term that scales with the batch. The step is launch-bound and the trace issues more launches than a plain one, so reach fortorch.compilebefore blaming plasticity for a slow run. - Ablate it; do not assume it. Temporal attention fills the same role on sequential classification and measured better there. The two are alternatives more often than complements.
# NLP / Logic / Reasoning — both mechanisms, for dynamic variable binding
model = OdyssNet(
num_neurons=32,
input_ids=[0, 1],
output_ids=[31],
activation='tanh',
hebb_type='both',
hebb_res='neuron', # Per-neuron plasticity
device='cuda',
)
# RL / reactive environments — neuron-level caste differentiation
model = OdyssNet(
num_neurons=64,
input_ids=list(range(8)),
output_ids=list(range(56, 64)),
activation='tanh',
hebb_type='temporal',
hebb_res='neuron', # Per-neuron plasticity
device='cuda',
)
# Quick experiment — global plasticity
model = OdyssNet(..., hebb_type='spatial', hebb_res='global')
# Default: ChaosGrad — estimates the step scale online, no tuning needed
trainer = OdyssNetTrainer(model)
# Fixed-rate mode: pass an explicit learning rate (reproducible)
trainer = OdyssNetTrainer(model, lr=3e-4)For tasks where the network must retrieve a specific past state rather than resonate with it — long-range recall, copying, variable binding over a long gap — enable attention over the core's own state history. There are no layers to stack it between, so it attends along time: every thinking step queries a cache of the states before it.
model = OdyssNet(
num_neurons=1024,
input_ids=list(range(256)),
output_ids=list(range(256, 768)),
vocab_size=2048, vocab_mode='discrete',
attn_heads=4, # the whole switch; None (default) builds nothing
attn_kv_heads=1, # multi-query: the cache, not the weights, is what OOMs
attn_window=256, # how far back a query can see
attn_write='token', # one entry per token, not per thinking step
attn_read='step', # 'token' halves the cost when think_gap > 0
device='cuda',
)- Free to try:
o_projis zero-initialized and the module is built after the core, so a model with attention and one without, at the same seed, have the sameWand produce the same output until training moves the attention weights. Ablations of it have one variable in them. - Widening it is safe: the branch's output is divided by
sqrt(attn_heads · attn_head_dim), so its reachable contribution does not grow with its width. Without that factor, a 4×64 branch collapsed associative recall to chance while a 1×16 branch matched the baseline — do not remove it, and if you add another parallel branch to the step, give it the same treatment. - What it costs: the query runs every step, so extra thinking steps multiply the reads, not the entries —
attn_read='token'is the knob for that. Keys and values kept for the backward pass scale as2·B·H_kv·D·(window + n²/2)withnthe writes inside one truncated-BPTT window;model.attn.training_cache_bytes(batch, n)returns the number, and the LLM harness prints it before training starts. - Cold starts: use
model.reset_rows(mask), never a masked fill onmodel.state. A row whose state is zeroed while its KV history survives is neither cold nor warm. - Optimizer: ChaosGrad classifies the four projections into an
attentionfamily (weight decay on), and the QK-norm gains intomodulation(decay off) — never decay those by hand in a custom optimizer. - Transplantation:
attn_head_dimdefaults to a value derived fromnum_neurons, so pin it explicitly when transplanting between cores of different sizes, or the attention geometry moves with the neuron count. - Compatibility: works with
gradient_checkpointing=True, AMP, neurogenesis and Hebbian plasticity; tested for all of them. - Precision: attention runs with autocast off, so the cache is single-dtype and the softmax accumulates in fp32; half precision measured ~0.017 of loss worse.
attend/writecarry@torch._dynamo.disablebecause Inductor miscompiles that region — do not remove it, every compiled run with attention on dies onHalf != float. The resulting graph break is free when the step has other work to fuse and costs the speedup when attention is the only thing switched on, so measure before assumingtorch.compilepays there. - Measure before you believe:
examples/advanced/experiment_llm.py --mode sweep --sweep attngives every arm the same wall-clock, withoffbuilding no attention at all. On TinyStories at 2 min/arm theoffarm wins — see docs/LIBRARY.md.
Always enable TF32 on Ampere+ GPUs for free speedup.
import torch
torch.set_float32_matmul_precision('high')The largest speedup this architecture has. The echo loop issues thousands of small kernels per step and is bound by launching them: eager → compiled is 2.2x on a bare core and 4.0x with plasticity.
model.compile() # or torch.compile(model.forward)Compile the bound forward rather than wrapping the module when other code
reaches through the model for state, KV caches, plastic buffers or
checkpointing — an OptimizedModule wrapper sits between them and the real
object.
Give it fixed shapes, or it will quietly stop compiling. Every new shape
costs a recompile, and enough of them exhaust Dynamo's budget and drop the
process back to eager for good, with nothing in the output saying so. Pass
drop_last=True to both loaders — a test set rarely divides by the batch size.
Experiments should handle training stagnation intelligently by adding neurons when needed.
- Metric: If
losshas not improved forNepochs. - Action: Call
trainer.expand(amount=...). - Amount:
- Small nets (< 100 neurons): +1 neuron per expansion.
- Large nets (≥ 100 neurons): +10 neurons or +1% of current size per expansion.
if loss > prev_loss:
trainer.expand(amount=10)Experiments that run for a long time should handle training stagnation or spikes intelligently without manual restarts.
You can pass an anomaly_hook to the OdyssNetTrainer to automate recovery and logging.
def my_hook(anomaly_type, loss_val):
if anomaly_type == "plateau":
print("Triggering plateau escape!")
trainer = OdyssNetTrainer(model, anomaly_hook=my_hook)Since tanh is our primary activation for robust systems, avoid using 0.0 and 1.0 for logical states.
- OFF:
-1.0 - ON:
1.0 - Neutral/Silence:
0.0
This symmetry helps the gradients flow much better than a 0.0 (which is the most unstable point of tanh).
Use the prepare_input utility implicitly via the Trainer.
- Pulse: Single Event at t=0.
- Stream: Continuous sequence. pass
full_sequence=Truetotrainer.predict()ortrain_batch()if you need frame-by-frame monitoring.
-
Reproducibility & Seeding: 🔴 MANDATORY — All example and experiment scripts MUST set a fixed seed for reproducible results.
- Why? Reproducible results are essential for debugging, comparing strategies, and publishing findings.
- How? Call
set_seed(...)at the very start of yourmain()function, before any random operations. Use 42 unless a measurement chose otherwise; the three MNIST record/tiny examples use 123. - Import:
from odyssnet import set_seed - Example:
def main(): set_seed(42) # ← FIRST LINE in main() # Now all randomness is locked: model init, data shuffling, dropout, etc. model = OdyssNet(...) trainer = OdyssNetTrainer(model, lr=1e-4) # pin lr for deterministic curves trainer.fit(X, Y, epochs=100)
- Applies to:
- Model weight initialization (deterministic via
torch.manual_seed). - Data shuffling / batch sampling.
- Dropout and stochastic regularization.
- CUDA random state (for GPU consistency).
- Model weight initialization (deterministic via
- Test: If you run the script twice with the same seed, loss curves and final results should be identical, byte-for-byte. Note: byte-for-byte identity requires ChaosGrad's fixed-rate mode (explicit
lr) — the defaultlr=Noneestimates the step scale online, so curves vary slightly across runs even when seeded.
-
Visuals: Your example should print a cool visualization. Don't just print "Loss: 0.01". Print the timeline.
- Example:
t=05 | Input: 1 | Output: 0.99 🟢
- Example:
-
Comments: Explain why you chose a specific setup.
- Example:
# GAP=3 allows the model time to digest the previous bit.
- Example:
-
File Paths: Never use hardcoded absolute paths or assume the CWD. Always construct paths relative to the script file.
- Example:
DATA_FILE = os.path.join(os.path.dirname(__file__), '..', '..', 'data', 'file.txt')
- Example:
-
Checkpointing: Always use the library's
save_checkpoint,load_checkpoint, andtransplant_weightsfunctions fromodyssnet.utils.odyssstore. Do NOT write custom checkpoint code. If the library is missing a feature, extend the library instead.- Example:
from odyssnet import save_checkpoint, load_checkpoint, transplant_weights
- Example:
-
Training History & Plotting: All finite-duration examples MUST use
TrainingHistoryto record metrics (loss, accuracy, lr, etc.) and callhistory.plot()at the end. This generates a multi-panel plot of all tracked metrics over time. (Note: When examples are run viatest_all.py, theODYSSNET_DISABLE_PLOT=1environment variable automatically bypasses the interactive plotting UI).- Example:
from odyssnet import TrainingHistory history = TrainingHistory() for epoch in range(epochs): loss = trainer.train_batch(...) history.record(loss=loss, lr=current_lr) history.plot(title="My Experiment")
-
Imports: Import
odyssnetdirectly (installed viapip install -e .). Never usesys.path.appendhacks.
If your experiment produces Loss nan / PPL nan, enable the built-in diagnosis mode before anything else:
model = OdyssNet(..., debug=True)With debug=True the model checks every critical forward-pass operation (linear recurrence, memory feedback, activation, StepNorm, Hebbian correlation and accumulation) and raises a RuntimeError at the first operation that produces a non-finite value, with the operation name and step index. debug=True also automatically enables torch.autograd.set_detect_anomaly(True), so backward-pass NaN is caught with a full stack trace at no extra setup cost. Overhead is zero when debug=False.
If your model trains but doesn't converge or gets stuck:
Track and visualize all key metrics to identify patterns:
from odyssnet import TrainingHistory
history = TrainingHistory()
for epoch in range(epochs):
loss = trainer.train_batch(x, y, thinking_steps=10)
acc = evaluate_accuracy(...)
lr = trainer.get_diagnostics()['current_lr'] # ChaosGrad's live estimate
history.record(loss=loss, accuracy=acc, lr=lr)
# Visual inspection reveals patterns
history.plot(title="Training Diagnosis")
# Or save for later analysis
history.plot(save_path="diagnosis/training.png", title="Debug Run")What to look for:
- Flat loss: May need more thinking steps, different initialization, or learning rate adjustment
- Oscillating loss: Reduce learning rate or enable gradient persistence
- Sudden spikes: Check for batch corruption or use anomaly_hook to catch them
Monitor training health in real-time:
for epoch in range(epochs):
loss = trainer.train_batch(x, y, thinking_steps=10)
# Get comprehensive diagnostics (add debug=True for detailed stats)
diag = trainer.get_diagnostics()
if epoch % 10 == 0:
print(f"Epoch {epoch}")
print(f" Step count: {diag['step_count']}")
print(f" Last loss: {diag['last_loss']:.6f}")
# For detailed debugging, use debug=True
if need_detailed_analysis:
diag = trainer.get_diagnostics(debug=True)
# Now includes gradient_stats, persistent_grads_active, anomaly_tracking,
# loss_tracking, scaler_state, and detailed optimizer per-param statsKey metrics to monitor:
- last_loss trend: Sustained increase suggests instability
- gradient_stats: Large norm swings can indicate optimization difficulty
- step_count progression: Confirms stable optimizer stepping cadence
Debug mode additions:
- gradient_stats: Per-parameter gradient norms and means (min/max/std)
- persistent_grads_active: Number of parameters with persistent gradients
- anomaly_tracking: EWMA, variance, and plateau detection state
- loss_tracking: Recent losses and buffer statistics
Set up intelligent automated responses to training anomalies:
def handle_anomaly(anomaly_type, loss_val):
"""Called automatically on training anomalies."""
if anomaly_type == "spike":
# Violent loss surge (possible gradient explosion)
print(f"⚠️ SPIKE detected! Loss: {loss_val:.4f}")
# Could reduce LR, reload checkpoint, etc.
elif anomaly_type == "plateau":
# Loss stagnated over a window
print(f"🔄 PLATEAU detected. Triggering escape...")
elif anomaly_type == "increase":
# Loss increased from previous step (happens every time loss goes up)
# Useful for custom patience counters or early stopping
global patience_counter
patience_counter += 1
if patience_counter > 50:
print(f"⛔ 50 consecutive increases. Early stopping.")
raise KeyboardInterrupt
# Initialize trainer with hook
patience_counter = 0
trainer = OdyssNetTrainer(
model,
anomaly_hook=handle_anomaly
)
# Now train — anomalies trigger automatic responses
for epoch in range(1000):
loss = trainer.train_batch(x, y, thinking_steps=10)Anomaly types:
- "spike": Sudden violent surge in loss (exploding gradient)
- "plateau": Loss stagnated and barely moving over a window
- "increase": Loss strictly greater than previous step (fired every time, even 0.0001 increase)
If loss oscillates or training is unstable:
-
Enable gradient persistence for smoother optimization:
trainer = OdyssNetTrainer(model, gradient_persistence=0.1)
-
Pin a lower explicit learning rate (ChaosGrad fixed-rate mode):
trainer = OdyssNetTrainer(model, lr=1e-4)
Or keep automatic estimation but make it more cautious:
from odyssnet import ChaosGrad trainer = OdyssNetTrainer(model, optimizer=ChaosGrad.from_model(model, d_coef=0.5))
-
Try different initialization if using tiny networks:
model = OdyssNet(..., weight_init='xavier_uniform', activation='gelu')
If loss doesn't decrease at all:
- Verify data preprocessing: Check that inputs/targets are properly normalized and on correct device
- Increase thinking steps: Model may need more temporal depth
trainer.train_batch(x, y, thinking_steps=20) # Was 10
- Check initialization: For very small networks (<10 neurons), try:
model = OdyssNet(..., weight_init='xavier_uniform', activation='gelu')
- Use anomaly_hook and adjust lr/steps dynamically based on diagnostics.
If training is too slow:
-
Enable TF32 on Ampere+ GPUs:
import torch torch.set_float32_matmul_precision('high')
-
Compile the model (PyTorch 2.0+):
model.compile()
-
Use gradient accumulation instead of larger batches:
# Simulates batch_size=128 with batch_size=32 trainer.train_batch(x, y, thinking_steps=10, gradient_accumulation_steps=4)
If running out of VRAM:
- Reduce batch size and use gradient accumulation
- Enable gradient checkpointing:
model = OdyssNet(..., gradient_checkpointing=True)
- Use vocab projection for high-dimensional inputs:
# Instead of num_neurons=784 for MNIST model = OdyssNet(num_neurons=10, vocab_size=[784, 10])
When modifying the library itself (not examples), follow these additional rules:
- Every change to a public interface must have a corresponding test update under
tests/. - The test suite mirrors
odyssnet/:tests/core/,tests/training/,tests/utils/. Place new tests in the matching subdirectory. - New behavior requires new test cases. Changed signatures require updated tests. Deleted code requires removed orphaned tests.
- Run
python -m pytest tests/or justpytest tests/(but not recommended as it will not see your env setup) after every change to confirm the suite stays green.
- Every public API change must be reflected in the relevant markdown files (
docs/LIBRARY.md,CONTRIBUTING.md). - Document what the system is, not what it was. Version history belongs in
CHANGELOG.mdonly.
- Comments explain why, not what — the code already says what.
- Use precise, professional language. Avoid filler phrases ("simply", "just", "obviously").
- Do not use comments or docs as changelog.
- Tests pass (
python -m pytest tests/orpytest tests/) - No
sys.path.appendhacks — imports usefrom odyssnet import ...directly
- Corresponding test added/updated under
tests/ - Documentation updated in relevant markdown files (docs/LIBRARY.md, CONTRIBUTING.md)
- Does your script call
set_seed(...)at the START ofmain()? (MANDATORY for reproducibility. Use 42 unless a measurement chose otherwise.) - Is the trainer zero-config (
OdyssNetTrainer(model, device=...)) unless there is a documented reason to pinlr? Zero-config ChaosGrad is the recommended default for examples. Pin an explicitlronly for precision-record scripts whose advertised metrics depend on a specific tuned rate — and say so in a comment. - Did you place it in the correct folder (
examples/for core validations,examples/advanced/for complex tasks)? - Are you using
OdyssNetTrainer? - Did you select the correct
activation,weight_init, andgatesetup? (Defaultresonant+gate=Noneis fine for most tasks.) - If you set
hebb_type, remember that ChaosGrad automatically classifies Hebbian logits into theplasticityfamily with zero weight decay — never add decay to them via a custom optimizer. 6b. [ ] If you setattn_heads, does the script reset rows withmodel.reset_rows(mask)rather than zeroingmodel.statedirectly, and does it say what the attention is measured against? An attention arm is only meaningful next to anattn_heads=Noneone at the same seed. - Does it converge reliably? (If you see
Loss nan, see Troubleshooting above.) - Does the terminal output clearly explain what is happening?
- Does the script use
TrainingHistoryand callhistory.plot()at the end? - Are file paths relative to
__file__, not hardcoded?
Welcome to the Order of the Algorithm. Let's code Time.