Impressive work!
For comparison, an H100 can do roughly 100-200 tok/s of Muse 30B for raw inference single-thread, going up to low thousands of tok/s with a large number of threads - and I am sure that for the massively-multi-threaded case they can optimize the prover further.