OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.
You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.
Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?
Yeah I read the title and did a double-take. It's an uninteresting choice for systems like these. The more pressing concerns, which go undescribed, are the exact mathematical choices behind the actual model. Rust provides almost zero value here because tensor stuff is all just 2d-arrays of floats for the most part.
You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.
Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?
Why did you choose Rust? Why does that matter?
Nothing wrong with Rust. Lots wrong with the bandwagon that “in Rust” somehow adds value.