Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

ย 

History

16 Commits
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

Chain-of-Thought (CoT) Collapse Benchmarks

An empirical evaluation framework designed to stress-test and map the exact logical decay, semantic degradation, and performance failure limits of large reasoning models under lossy token compression and multi-step inference execution.


๐Ÿ“ Evaluation Framework

  • Deterministic Tracking: Measuring token-to-logic degradation metrics.
  • Semantic Compression Analysis: Testing inference stability when prompt structures are highly compressed.
  • Failure Layer Mapping: Identifying the exact depth (step count) where computational reasoning collapses.

๐ŸŽ›๏ธ Pipeline Engine

  • Core Engine: Python 3.x
  • Architecture: Autonomous JSON automation pipelines for prompt stress-testing.

Status: Architecture active. Core mathematical benchmarking scripts deploying next..

About

Empirical evaluation of Chain-of-Thought reasoning degradation under lossy semantic compression.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages