Rexium Forge

Specialised AI that runs where your data is.

Ingot is our compressed base model. It fits in 1.45 GB, runs on a laptop or a small server, and answers without sending anything outside your network.

1.45 GB on diskApache-2.0European Portuguese and English

The AI that impresses lives in a datacentre.

It needs a permanent connection, charges for every use, and requires sending your data to someone else's servers. For a clinic, a gym, a factory or anywhere without reliable coverage, that is expensive, risky or simply impossible. Small models that run on cheap hardware do exist, but they are generic: they know a little about everything and little about what your business actually does.

Small by design, not by cutting.

Shrinking a model is easy. Shrinking one without taking away what made it useful is not, and a badly shrunk model gives itself away by the third answer. Ingot was compressed with the method we developed in the Rexium Forge, and what matters to whoever uses it is the result: it fits in 1.45 GB and still answers at the level of a model several times its size. The numbers are right below, caveat in plain sight.

The result

In our internal exam, Ingot was indistinguishable from a model twice its size while using five times less memory.

The numbers, with the caveat in plain sight

All measured on the same GPU, with the same frozen exam and reference answers from the domain expert.

ModelScoreMemory
Qwen3-4B (uncompressed)51,7%7,50 GB
Ingot-2B (6-bit)51,7%1,44 GB
Qwen3.5-2B (uncompressed)47,9%3,76 GB

What this does not mean

With a 12-item exam and judge noise measured at about 1.5 points, saying two models tie means we could not separate them, not that they are identical. This is an internal exam, not a public benchmark. We publish it this way because we prefer an honest number to a pretty one, and because it is what any serious buyer will ask.

One base. Many specialists.

Ingot is the raw material, not the finished product. Each business trains on top of it with what it knows, and ends up with a model that understands its domain without ever having shared the data with anyone.

Ingot-2B

The base. Compressed once, the same for everyone, public on request.

Ingot-2B-sport

The first specialist, in progress: high-performance training, from a domain professor's own material, to power the PeakRaptor platform.

Yours

What your business knows and nobody else has. The specialist weights are yours and stay private.

Who this is for

Where the network fails

Basement gyms, factories, construction sites, fieldwork. Places where the connection drops and the work cannot wait for it.

Where data cannot leave

Clinics, health data, data about minors, industrial information. The model travels to the data instead of the data travelling to the model.

Where per-use pricing doesn't add up

High volume, thin margins. A model running on hardware you already own doesn't send an invoice at the end of the month.

Technical sheet

Base model
Qwen3.5-2B
Licence
Apache-2.0, no revenue cap, no use restrictions
Format
GGUF, Q6_K quantisation (6-bit)
Size
about 1.45 GB on disk
Languages
European Portuguese and English
Access
Public on Hugging Face, with manual approval

Ingot derives from Qwen3.5-2B, released under Apache-2.0. What is ours is the compression and the training on top, not the pre-training.

Want a specialist of your own?

Tell us what your business knows and where it needs to run. We reply within 24 hours, and the first reply is from a person.