-
Notifications
You must be signed in to change notification settings - Fork 83
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Question about performance comparison #323
Comments
Same question here: any comparison with oneDNN ? |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Hi.
I was looking for a performance comparison between
ruy
andOpenBLAS
and I came across this.But when I benchmark the ruy (almost for any shape with single thread execution and on raspberry pi 4), my results are far behind the reported results.
For example, for the 512x512x512 Int8 benchmark, I can only get ~10 GOPs but excel reported 40 GOPs.
I know Raspberry Pi 4 CPU frequency can be maxed out to 1.5 GHz while Pixel 4 max frequency is 2.84 GHz, but it does not justify the 30 GOPs gap.
So I thought it might be better to ask it here.
How did you measure GOPs for ruy?
I calculate the GOPs for the method with the
((2 * N * K * M * iterations) / time) / 10e+9
formula (time
is the sum of the execution time ofruy::Mul
for each iteration) (I pack the RHS matrix beforehand).Am I doing anything wrong?
The text was updated successfully, but these errors were encountered: