Respan Announces Span-01: The First Hyper-Parallel Reasoning Classifier Built for Unseen Challenges


Respan, defining the next generation of the agent infra stack, led by Andy Li, Raymond Huang, and Hendrix Liu, has announced Span-01: the first hyper-parallel reasoning classifier built for unseen challenges.
Jev demonstrated that decision models can make classification fast and practical. The next challenge is whether a classifier can preserve frontier-level reasoning, generalize across unseen definitions, and carry that reasoning through long traces in a single forward pass. Span-01 is designed to reason over new natural-language definitions and return direct probabilities across every definition in parallel.
Reason First. Classify once.
The difficult part of classification is not returning a label. It is applying a definition never seen before to context never encountered before.
Span-01 is trained with general classification reasoning through RLAIF before being specialized for behavior detection. Rather than memorizing a fixed taxonomy, the model executes unseen behavior definitions across new behaviors and domains.
At inference, hyper-parallel definition branches evaluate multiple unseen behaviors over the same trace context. Hybrid attention helps preserve that reasoning across long, complex traces.
One true forward pass. No token-by-token generation. Span-01 returns direct probabilities for present, absent, and not_observable.
Frontier-Level Judgment.
The first test was whether a dedicated classifier could retain the judgment quality of frontier models. Across nine models and both English and multilingual behavior detection, Span-01 ranks first overall. Span-01 Lite also outperforms Jev and Sonnet 5.
Built for the Behaviors Production Teams Actually Monitor.
Aggregate scores can obscure the behaviors that matter in production. The evaluation separates safety, reliability, grounding, agent behavior, and response quality to identify where the signal holds. Span-01 outperforms Jev and Sonnet 5 overall, while a full frontier model remains the strongest performance ceiling.
Decision Models Judged by Span-01.
The benchmark was then inverted, with Span-01 serving as the evaluation layer for 11 emerging decision models across accuracy, consistency, injection resistance, and calibration. Span-01 does not appear in this ranking. It produced the evaluation signal.
𝐑𝐞𝐦𝐨𝐭𝐞 𝐇𝐢𝐫𝐞 𝐨𝐫 𝐁𝐮𝐢𝐥𝐝 𝐖𝐢𝐭𝐡 𝐔𝐬: 𝐁𝐮𝐢𝐥𝐝𝐰𝐞𝐫𝐤𝐬 helps technology companies build in India: two ways. Hire vetted remote engineers, product leaders, designers, and GTM professionals directly onto your team. Or partner with a proven development studio to design, build, and ship your product end-to-end. Learn More At: menlotimes.com/buildwerks


