Department
of Computer Science
School of Engineering and
Applied Science
University of Virginia, Charlottesville,
Virginia
STREAM: Sustainable Memory Bandwidth in High Performance Computers
Table of Contents
-
What's New?
-
Version 5 of the STREAM benchmark is available in Fortran! Check
it out.
-
STREAM FAQ has been significantly revised
and enhanced, with more discussion of how to run the STREAM benchmark.
Check it out!
-
The Archives of Original Submissions of STREAM benchmark results have been
updated through April 18, 2000! All of the results have made it into the
tables, but the links from the tables back to the original submissions
are still being (slowly & manually) added.
-
Here are the current RESULTS!
-
STREAM FAQ
-
Analyses, Commentary, etc....
-
Hypermail archives of original contributions
-
FTP access to code and data:
Many people have reported that the anonymous ftp service is often unavailable,
especially on weekends. I do not know what is happening, but the service
does get restarted, and seems to be up more often during the work week,
so keep on trying!
Here are the current RESULTS!
Top 10 Results for Shared-Memory Systems!
This set of results includes the top 10 shared-memory systems, ranked by
STREAM TRIAD performance. The results are currently presented in the following
tables:
Standard Results
The "standard" set of results presents the results of the C or Fortran
versions of the STREAM benchmark running with 64-bit data types on production
hardware. The "standard" set of results excludes:
-
Cases with 32-bit operands or operations.
-
Cases with significantly modified code or assembly language.
-
Simulated results.
-
Results from experimental or non-production hardware or software.
The results are currently presented in the following tables:
PC-compatible Results
This set of tables summarizes the "standard" test cases, but restricted
to IBM PC-compatible computers.
Users are free to re-compile the code
or use the new NT
or Linux
binaries. Use of the old DOS binaries is discouraged.
The results are currently presented in the following tables:
Macintosh-Compatible Results
This set of tables summarizes the "standard" test cases, but restricted
to Macintosh and compatible computers.
Users are free to re-compile the code
or use the contributed
binaries.
The results are currently presented in the following tables:
Experimental/Nonstandard Results
These tables include only results that are
-
Assembly language coded, or
-
Simulated, or
-
Based on experimental or non-production hardware, or
-
Based on partially depopulated systems.
Note:
A "partially depopulated" system is one in which only a subset of the
cpus are used for the benchmark, and for which this subset is spread around
the machine to decrease contention. For example on the SGI Origin2000,
each node has 2 cpus sharing a single bus and memory subsystem. The results
in this table labelled "1 per node" are based on using only one cpu per
node board, and are considered a "nonstandard" way of using the machine.
Similarly, the Sun Ultra10000 has 4 cpus per node board, so results using
1, 2, or 3 cpus per node also go into this table of "nonstandard" results.
The results are currently presented in the following tables:
32-bit Results
These tables include only results with 32-bit operands or operations.
Results using 64-bit operands that move the data in 32-bit "chunks" are
not
here. The results are currently presented in the following tables:
John D. McCalpin mccalpin@cs.virginia.edu