CIR
Cost efficient intelligence research.
CIR investigates architectures and learning systems that may lower total cost. The aim is equivalent model capability for less economic cost.
CIR is active research. No benchmark results have been published yet. Current experimental architectures are not the final CIR architecture.
Can a different architecture reach the same capability for substantially less total cost?
The comparison point is a strong Transformer system. Q stands for matched useful model capability.
CIR searches for designs that make the ratio below meaningfully smaller than one.
minimize TotalCost(CIR, Q) / TotalCost(Strong Transformer, Q)
A ratio of 0.5 would mean equal capability at half the total cost. This is the objective, not a result.
Proxies are useful. None of them is the cost.
CIR does not optimize any single proxy on its own. It asks a more economically meaningful question. How much total cost does a given level of capability require?
FLOPs
Ignores memory traffic and utilizationToken throughput
Says little about capability reachedParameter count
Active compute can differ greatlyBits per byte
One view of capability, not all of itCPU time
Depends on hardware and implementationInference speed
Leaves out training cost entirely
How CIR evaluates a candidate.
Why the Transformer
It is the strongest, most studied and most optimized design available. Beating a weak baseline proves little.
Why BPB alone is insufficient
Bits per byte measures compression of text. Two models with equal BPB can differ in recall, reasoning and long context use.
Capability is multidimensional
Candidates are compared across several capability dimensions. A gain on one dimension cannot hide a loss on another.
Hardware aware evaluation
Cost is measured on real CPUs and GPUs. A design that saves FLOPs but stalls on memory can cost more.
Matched and fair comparisons
Baselines get the same tuning effort and budget. Results are checked across seeds and scales before any claim.
Measured, estimated, projected
These three kinds of numbers stay clearly labeled. Unexpected failures are recorded, not hidden.
Where CIR looks for cost.
01Neural architecture
Model structure02Recurrent and state mechanisms
Model structure03Attention
Model structure04Memory
Model structure05Learning efficiency
Learning06Optimizer interaction
Learning07Parameter activity
Learning08Data efficiency
Learning09CPU and GPU hardware behavior
Systems10Memory traffic
Systems11Training systems
Systems12Capability evaluation
Evaluation
What CIR does not claim.
CIR is part of DotrixAI, not the whole lab. Other programs may follow.
It has not replaced the Transformer.
It is not proven at frontier scale.
It claims no fixed cost multiple over Transformers.
It claims no proven reasoning advantage.
It makes no claim of outside adoption.
Today's experimental designs are not the final architecture.