From kaul@ee.eng.ohio-state.edu  Fri Sep 27 21:54:21 1991
Received: from everest-o.eng.ohio-state.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA16278; Fri, 27 Sep 91 21:54:21 EDT
Received: from nile (nile.eng.ohio-state.edu) by everest (4.1/3.910501)
	id AA28110; Fri, 27 Sep 91 21:53:16 EDT
From: kaul@ee.eng.ohio-state.edu (Rich Kaul)
Received: by nile (4.1/EMI-2.0)
	id AA02319; Fri, 27 Sep 91 21:51:34 EDT
Date: Fri, 27 Sep 91 21:51:34 EDT
Message-Id: <9109280151.AA02319@nile.eng.ohio-state.edu>
To: mccalpin (John D. McCalpin)
In-Reply-To: mccalpin@perelandra.cms.udel.edu's message of 28 Sep 91 01:23:26 GMT
Subject: Attainable memory bandwidth
Status: RO

   I am ***very*** interested in results from the following:
	   HP 7x0 machines
	   SGI 4D/35
	   SGI Indigo
	   Sun SPARC 2

For a Sparc 2:

Timing calibration ; t =     49.0000 clicks
    
Assignment: Rate =     32.6531 MB/s MFLOPS =   0.
Scaling:    Rate =     27.5862 MB/s MFLOPS =     1.72414
Summing:    Rate =     35.8209 MB/s MFLOPS =     1.49254
SAXPYing:   Rate =     24.4898 MB/s MFLOPS =     2.04082

-rich

From broadley@schneider3.lrdc.pitt.edu  Fri Sep 27 22:29:43 1991
Received: from schneider3.lrdc.pitt.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA16287; Fri, 27 Sep 91 22:29:43 EDT
Received: by schneider3.lrdc.pitt.edu (5.57/Ultrix3.0-C)
	id AA19347; Fri, 27 Sep 91 22:28:04 -0400
Date: Fri, 27 Sep 91 22:28:04 -0400
From: broadley@schneider3.lrdc.pitt.edu (Bill Broadley)
Message-Id: <9109280228.AA19347@schneider3.lrdc.pitt.edu>
To: mccalpin (John D. McCalpin)
Subject: Re: Attainable memory bandwidth
Status: RO


Heres the output for a SparcII:
	Unfortunately xlock, and xnews are running in the background.
	
mthbard% a.out
Timing calibration ; t =    104.0000 clicks

Assignment: Rate =     15.3846 MB/s MFLOPS =   0.
Scaling:    Rate =     15.8416 MB/s MFLOPS =    0.990099
Summing:    Rate =     19.0476 MB/s MFLOPS =    0.793651
SAXPYing:   Rate =     15.5844 MB/s MFLOPS =     1.29870
mthbard%

	I am getting a 720 Real soon now, I will let you know.
	Please post a new summary.
			-Bill
			Broadley@schneider3.lrdc.pitt.edu

From uunet.UU.NET!ssigv!lewis  Fri Sep 27 22:56:54 1991
Received: from relay2.UU.NET by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA16293; Fri, 27 Sep 91 22:56:54 EDT
Received: from uunet.uu.net (via LOCALHOST.UU.NET) by relay2.UU.NET with SMTP 
	(5.61/UUNET-internet-primary) id AA20366; Fri, 27 Sep 91 22:55:24 -0400
Received: from sunrise.UUCP by uunet.uu.net with UUCP/RMAIL
	(queueing-rmail) id 225354.22749; Fri, 27 Sep 1991 22:53:54 EDT
Received: by gv.ssi.com (4.1/SMI-4.1)
	id AA10009; Fri, 27 Sep 91 19:44:26 PDT
Date: Fri, 27 Sep 91 19:44:26 PDT
From: uunet.UU.NET!ssigv!lewis (Don Lewis)
Message-Id: <9109280244.AA10009@gv.ssi.com>
To: mccalpin
Subject: Re: Attainable memory bandwidth
Newsgroups: comp.benchmarks,comp.arch
In-Reply-To: <MCCALPIN.91Sep27212326@pereland.cms.udel.edu>
Organization: Silicon Systems, Nevada City CA
Cc: 
Status: RO

Sun SPARC 2

f77 version 1.3.1 with no options:

Timing calibration ; t =     64.0000 clicks
    
Assignment: Rate =     25.0000 MB/s MFLOPS =   0.
Scaling:    Rate =     15.5340 MB/s MFLOPS =    0.970874
Summing:    Rate =     26.9663 MB/s MFLOPS =     1.12360
SAXPYing:   Rate =     18.3206 MB/s MFLOPS =     1.52672


f77 version 1.3.1 with "-cg89 -dalign -libmil -O3" options

Timing calibration ; t =     36.0000 clicks
    
Assignment: Rate =     44.4445 MB/s MFLOPS =   0.
Scaling:    Rate =     21.6216 MB/s MFLOPS =     1.35135
Summing:    Rate =     45.2830 MB/s MFLOPS =     1.88679
SAXPYing:   Rate =     23.5294 MB/s MFLOPS =     1.96078

-- 
Don "Truck" Lewis              Phone: +1 916 265-3211   Silicon Systems
Internet: (under contruction)  FAX:   +1 916 265-2931   138 New Mohawk Road
UUCP: {uunet,tektronix!gvgpsa.gvg.tek.com}!ssigv!lewis  Nevada City, CA  95959


From alz@grumpy.cray.com  Sat Sep 28 08:56:43 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA16745; Sat, 28 Sep 91 08:56:43 EDT
Received: from grumpy.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6j)
	id AA06897; Sat, 28 Sep 91 07:54:55 CDT
Received: by grumpy.cray.com
	id AA13416; 4.1/CRI-4.4; Sat, 28 Sep 91 08:54:52 EDT
Date: Sat, 28 Sep 91 08:54:52 EDT
From: alz@grumpy.cray.com (Andrew Zachary)
Message-Id: <9109281254.AA13416@grumpy.cray.com>
To: mccalpin
Subject: Cray Results for bandwidth test
Cc: alz@grumpy.cray.com
Status: RO


John,

I got curious, so I took your benchmark test and ran it on a Y-MP.

I used sn1057, an 8 processor, 128 Mword Y-MP with 256 memory banks.
Here are the results from your test....

 Timing calibration ; t = 0.1128156 clicks
     
 Assignment: Rate = 14182.43576243 MB/s MFLOPS = 0.
 Scaling:    Rate = 14185.52898724 MB/s MFLOPS = 886.5955617026
 Summing:    Rate = 21278.29348086 MB/s MFLOPS = 886.5955617026
 SAXPYing:   Rate = 21277.27480664 MB/s MFLOPS = 1773.106233887

From olson@anchor.esd.sgi.com  Sun Sep 29 00:12:09 1991
Received: from SGI.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17047; Sun, 29 Sep 91 00:12:09 EDT
Received: from [192.26.58.38] by sgi.sgi.com via SMTP (910911.SGI.EXPERIMENTAL/910110.SGI)
	for mccalpin@perelandra.cms.udel.edu id AA14163; Sat, 28 Sep 91 21:10:29 -0700
Received: by anchor.esd.sgi.com (910711.SGI/910805.SGI)
	for @sgi.sgi.com:mccalpin@perelandra.cms.udel.edu id AA19764; Sat, 28 Sep 91 21:10:25 -0700
Date: Sat, 28 Sep 91 21:10:25 -0700
From: olson@anchor.esd.sgi.com (Dave Olson)
Message-Id: <9109290410.AA19764@anchor.esd.sgi.com>
To: mccalpin
Subject: Re: Attainable memory bandwidth
Newsgroups: comp.benchmarks,comp.arch
References: <MCCALPIN.91Sep27212326@pereland.cms.udel.edu>
Status: RO

In comp.arch you write:

| ==========================================================
| Memory Transfer Rates in MB/s for Standard Fortran Kernels
| ==========================================================
| John D. McCalpin			September 24, 1991
| mccalpin@perelandra.cms.udel.edu	DELOCN::MCCALPIN
| 
| The following table presents the results of a simple test of the
| memory bandwidth of a computer running some simple vector kernels
| coded in Fortran.  In each case, only the memory transfer rate in
| Millions of Bytes per second (MB/s) is shown.
| 
| Machine			Copy	Scale	Sum	SAXPY
| ------------------------------------------------------
| IBM RS/6000-950		190.5	163.3	184.6	187.5
| IBM RS/6000-530		145.5    88.9   109.1   120.0
| IBM RS/6000-320		 58.2	 55.7	 60.0	 60.0
| SGI 4D/240		 35.6	 22.1	 34.9	 25.6
| DEC 5000		 25.9	 25.6	 23.1	 21.8
| Sun 4/490		 25.0	 16.7	 24.0	 19.4
| Sun SS1+		 28.6	 12.3	 25.5	 13.3
| SGI 4D/25		 23.7	 11.6	 17.8	 13.7
| Sun SS1			 24.6	  9.3	 24.0	  9.6

Here are the results for an Indigo running 4.0, and a 4D/35,
also running the released IRIX 4.0.  I compiled it as f77 -O.  I'm not quite sure
that I believe the accuracy of the timing, since it runs so fast.
Both machines were multiuser, but idle.  Indigo should be somewhat
slower (slower clock, same memory system, smaller cache).
but Assignment doesn't seem to follow this model.

% uname -a
IRIX oceana 4.0 08281003 IP12
% hinv
1 33 MHZ IP12 Processor
FPU: MIPS R2010A/R3010 VLSI Floating Point Chip Revision: 4.0
CPU: MIPS R2000A/R3000 Processor Chip Revision: 3.0
On-board serial ports: 2
Data cache size: 32 Kbytes
Instruction cache size: 32 Kbytes
Main memory size: 24 Mbytes
Integral Ethernet: ec0, version 0
Disk drive: unit 1 on SCSI controller 0
Integral SCSI controller 0: Version WD33C93A
Iris Audio Processor, rev 2
Graphics board: LG1
% repeat 3 /tmp/X
 Timing calibration ; t =    29.00000     clicks
     
 Assignment: Rate =    55.17242     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    29.09091     MB/s MFLOPS =    1.818182    
 Summing:    Rate =    66.66664     MB/s MFLOPS =    2.777777    
 SAXPYing:   Rate =    47.05882     MB/s MFLOPS =    3.921569    
 Timing calibration ; t =    16.00000     clicks
     
 Assignment: Rate =    100.0000     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    32.65306     MB/s MFLOPS =    2.040816    
 Summing:    Rate =    74.99998     MB/s MFLOPS =    3.125000    
 SAXPYing:   Rate =    42.85715     MB/s MFLOPS =    3.571429    
 Timing calibration ; t =    16.00000     clicks
     
 Assignment: Rate =    100.0000     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    33.33333     MB/s MFLOPS =    2.083333    
 Summing:    Rate =    57.14286     MB/s MFLOPS =    2.380953    
 SAXPYing:   Rate =    47.05882     MB/s MFLOPS =    3.921569    



Here are the results for a 4D/35:

% uname -a
IRIX mears 4.0 08281003 IP12
% hinv
1 36 MHZ IP12 Processor
FPU: MIPS R2010A/R3010 VLSI Floating Point Chip Revision: 4.0
CPU: MIPS R2000A/R3000 Processor Chip Revision: 3.0
On-board serial ports: 4
Data cache size: 64 Kbytes
Instruction cache size: 64 Kbytes
Main memory size: 16 Mbytes
Integral Ethernet: ec0, version 0
Disk drive: unit 2 on SCSI controller 0
Disk drive: unit 1 on SCSI controller 0
Integral SCSI controller 0: Version WD33C93A
Graphics board: GR1.2 Bit-plane, Z-buffer, Turbo options installed
% repeat 3 /tmp/X
repeat 3 /tmp/X
     
 Assignment: Rate =    48.48484     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    30.76923     MB/s MFLOPS =    1.923077    
 Summing:    Rate =    68.57143     MB/s MFLOPS =    2.857143    
 SAXPYing:   Rate =    47.05882     MB/s MFLOPS =    3.921569    
 Timing calibration ; t =    21.00000     clicks
     
 Assignment: Rate =    76.19046     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    31.37255     MB/s MFLOPS =    1.960784    
 Summing:    Rate =    68.57145     MB/s MFLOPS =    2.857144    
 SAXPYing:   Rate =    46.15383     MB/s MFLOPS =    3.846152    
 Timing calibration ; t =    16.00000     clicks
     
 Assignment: Rate =    100.0000     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    33.33333     MB/s MFLOPS =    2.083333    
 Summing:    Rate =    74.99998     MB/s MFLOPS =    3.125000    
 SAXPYing:   Rate =    46.15385     MB/s MFLOPS =    3.846154    



From brooks@tazdevil.llnl.gov  Sun Sep 29 13:23:05 1991
Received: from tazdevil.llnl.gov by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17474; Sun, 29 Sep 91 13:23:05 EDT
Received: by tazdevil.llnl.gov (4.1/1.15)
	id AA28197; Sun, 29 Sep 91 10:21:39 PDT
Date: Sun, 29 Sep 91 10:21:39 PDT
From: brooks@tazdevil.llnl.gov (Eugene D. Brooks III)
Message-Id: <9109291721.AA28197@tazdevil.llnl.gov>
To: mccalpin
Subject: your kernels on a SPARC-2
Status: RO


Here are your kernels, done on a SPARC-2.  The run was done several
time to check for noise.

Timing calibration ; t =     42.0000 clicks
    
Assignment: Rate =     38.0952 MB/s MFLOPS =   0.
Scaling:    Rate =     20.5128 MB/s MFLOPS =     1.28205
Summing:    Rate =     42.1053 MB/s MFLOPS =     1.75439
SAXPYing:   Rate =     23.5294 MB/s MFLOPS =     1.96078
Timing calibration ; t =     41.0000 clicks
    
Assignment: Rate =     39.0244 MB/s MFLOPS =   0.
Scaling:    Rate =     20.7792 MB/s MFLOPS =     1.29870
Summing:    Rate =     41.3793 MB/s MFLOPS =     1.72414
SAXPYing:   Rate =     23.3010 MB/s MFLOPS =     1.94175
Timing calibration ; t =     42.0000 clicks
    
Assignment: Rate =     38.0952 MB/s MFLOPS =   0.
Scaling:    Rate =     20.7792 MB/s MFLOPS =     1.29870
Summing:    Rate =     41.3793 MB/s MFLOPS =     1.72414
SAXPYing:   Rate =     23.3010 MB/s MFLOPS =     1.94175
Timing calibration ; t =     42.0000 clicks
    
Assignment: Rate =     38.0953 MB/s MFLOPS =   0.
Scaling:    Rate =     20.7792 MB/s MFLOPS =     1.29870
Summing:    Rate =     41.3793 MB/s MFLOPS =     1.72414
SAXPYing:   Rate =     21.6216 MB/s MFLOPS =     1.80180
Timing calibration ; t =     44.0000 clicks
    
Assignment: Rate =     36.3636 MB/s MFLOPS =   0.
Scaling:    Rate =     20.7792 MB/s MFLOPS =     1.29870
Summing:    Rate =     41.3793 MB/s MFLOPS =     1.72414
SAXPYing:   Rate =     23.3010 MB/s MFLOPS =     1.94175

From alz@grumpy.cray.com  Sun Sep 29 13:34:33 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17489; Sun, 29 Sep 91 13:34:33 EDT
Received: from ivan.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6j)
	id AA26432; Sun, 29 Sep 91 12:32:55 CDT
Received: by ivan.cray.com
	id AA00367; 4.1/CRI-4.4; Sun, 29 Sep 91 13:32:50 EDT
Date: Sun, 29 Sep 91 13:32:50 EDT
From: alz@grumpy.cray.com (Andrew Zachary)
Message-Id: <9109291732.AA00367@ivan.cray.com>
To: mccalpin
Subject: Re:  Cray Results for bandwidth test
Cc: alz@grumpy.cray.com
Status: RO

John,

For the test results I sent you, the compiler options were

	cf77 -Zv ....

i.e., I did not explicitly invoke the multitasking libraries.  Now,
compiler and fpp may recognize SAXPY's and replace them
with the equivalent BLAS level 1 routines, which may be multitasked.
Since I did not run with   setenv NCPUS 1, the multitasked routines
would have run on all 8 CPU's.  However, I am skeptical, because
I ran your test several times, and each time I got the same
answers.  Could there be a subtle error in your test?

As to the granularity of second, it should be good down to the
clock period - 6 nsec.  Again, the answers differed only slightly
from run to run.

On Monday morning, when more machines are up, I'll try runnin your
test on a variety of Y-MP configurations.

Andrew Zachary
P.S.  I'll also increase the loop length to 5,000,000 or so...



From Keith.Bierman@Eng.Sun.COM  Sun Sep 29 15:19:41 1991
Received: from sun.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17513; Sun, 29 Sep 91 15:19:41 EDT
Received: from Eng.Sun.COM (zigzag-bb.Corp.Sun.COM) by Sun.COM (4.1/SMI-4.1)
	id AA18608; Sun, 29 Sep 91 12:17:46 PDT
Received: from chiba.Eng.Sun.COM by Eng.Sun.COM (4.1/SMI-4.1)
	id AA25870; Sun, 29 Sep 91 12:17:37 PDT
Received: by chiba.Eng.Sun.COM (4.1/SMI-4.1)
	id AA00243; Sun, 29 Sep 91 12:17:28 PDT
Date: Sun, 29 Sep 91 12:17:28 PDT
From: Keith.Bierman@Eng.Sun.COM (Keith Bierman fpgroup)
Message-Id: <9109291917.AA00243@chiba.Eng.Sun.COM>
To: mccalpin (John D. McCalpin)
In-Reply-To: mccalpin@perelandra.cms.udel.edu's message of 28 Sep 91 01:23:26 GMT
Subject: Attainable memory bandwidth
Status: RO


4/75 times.

32mb ram, sunos 4.1.1 (technically a post beta thereof; I never got around
to upgrading, and also a downrev prerelease prom .... but neither of
those things are known to affect performance). f77v1.4 patch 2 code
compiled -fast -O4 -Bstatic.

Five runs, triggered by

	for i in 1 2 3 4 5; do echo $i; a.out >> mc.out ; done

and running sans window system. I wasn't inclined to disconnect my
machine from the net; but I don't think anyone was mounting my
exported volumes and performing useful work during this test.

Timing calibration ; t =     44.0000 clicks
    
Assignment: Rate =     36.3636 MB/s MFLOPS =   0.
Scaling:    Rate =     29.6296 MB/s MFLOPS =     1.85185
Summing:    Rate =     38.0952 MB/s MFLOPS =     1.58730
SAXPYing:   Rate =     25.2632 MB/s MFLOPS =     2.10526
Timing calibration ; t =     44.0000 clicks
    
Assignment: Rate =     36.3636 MB/s MFLOPS =   0.
Scaling:    Rate =     29.6296 MB/s MFLOPS =     1.85185
Summing:    Rate =     38.0952 MB/s MFLOPS =     1.58730
SAXPYing:   Rate =     25.2632 MB/s MFLOPS =     2.10526
Timing calibration ; t =     44.0000 clicks
    
Assignment: Rate =     36.3636 MB/s MFLOPS =   0.
Scaling:    Rate =     29.6296 MB/s MFLOPS =     1.85185
Summing:    Rate =     38.0953 MB/s MFLOPS =     1.58730
SAXPYing:   Rate =     25.2631 MB/s MFLOPS =     2.10526
Timing calibration ; t =     44.0000 clicks
    
Assignment: Rate =     36.3637 MB/s MFLOPS =   0.
Scaling:    Rate =     29.6296 MB/s MFLOPS =     1.85185
Summing:    Rate =     38.0952 MB/s MFLOPS =     1.58730
SAXPYing:   Rate =     25.2632 MB/s MFLOPS =     2.10526
Timing calibration ; t =     44.0000 clicks
    
Assignment: Rate =     36.3636 MB/s MFLOPS =   0.
Scaling:    Rate =     29.6296 MB/s MFLOPS =     1.85185
Summing:    Rate =     38.7097 MB/s MFLOPS =     1.61290
SAXPYing:   Rate =     25.2632 MB/s MFLOPS =     2.10526


totals (arithmetic mean, standard deviation):

Assignment: avg = 36.3636  stdev =   4.47214e-05 #pts   5  
Scaling   : avg = 29.6296  stdev =   0 #pts   5   
Summing   : avg = 38.2181  stdev =   0.274802 #pts   5
SAXPYing  : avg = 25.2632  stdev =   4.47214e-05 #pts   5

cheers

From jsw@xhead.esd.sgi.com  Sun Sep 29 19:30:22 1991
Received: from SGI.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17623; Sun, 29 Sep 91 19:30:22 EDT
Received: from [192.48.198.11] by sgi.sgi.com via SMTP (910911.SGI.EXPERIMENTAL/910110.SGI)
	for mccalpin@perelandra.cms.udel.edu id AA03435; Sun, 29 Sep 91 16:28:42 -0700
Received: by xhead.esd.sgi.com (910711.SGI/910709.SGI)
	for @sgi.sgi.com:mccalpin@perelandra.cms.udel.edu id AA07783; Sun, 29 Sep 91 16:28:41 -0700
Date: Sun, 29 Sep 91 16:28:41 -0700
From: jsw@xhead.esd.sgi.com (Jeff Weinstein)
Message-Id: <9109292328.AA07783@xhead.esd.sgi.com>
To: mccalpin (John D. McCalpin)
Subject: Re: Attainable memory bandwidth
Status: RO

Compiled f77 -O2

SGI 4D/35:

 Timing calibration ; t =    16.00000     clicks
     
 Assignment: Rate =    100.0000     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    33.33333     MB/s MFLOPS =    2.083333    
 Summing:    Rate =    55.81396     MB/s MFLOPS =    2.325582    
 SAXPYing:   Rate =    47.05882     MB/s MFLOPS =    3.921569    


SGI Indigo:

 Timing calibration ; t =    24.00000     clicks
     
 Assignment: Rate =    66.66666     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    27.58620     MB/s MFLOPS =    1.724138    
 Summing:    Rate =    58.53657     MB/s MFLOPS =    2.439024    
 SAXPYing:   Rate =    38.09525     MB/s MFLOPS =    3.174604    

Will you be posting another set of results when you get some replys?

	--Jeff

-- 
Jeff Weinstein - X Protocol Police
Silicon Graphics, Inc., Entry Systems Division, Window Systems
jsw@xhead.esd.sgi.com
Any opinions expressed above are mine, not sgi's.


From blythe@banshee.asd.sgi.com  Sun Sep 29 20:22:17 1991
Received: from SGI.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17636; Sun, 29 Sep 91 20:22:17 EDT
Received: from giraffe.asd.sgi.com by sgi.sgi.com via SMTP (910911.SGI.EXPERIMENTAL/910110.SGI)
	for mccalpin@perelandra.cms.udel.edu id AA04650; Sun, 29 Sep 91 17:20:41 -0700
Received: from banshee.asd.sgi.com by giraffe.asd.sgi.com (5.52/900721.SGI)
	for sgi.sgi.com!perelandra.cms.udel.edu!mccalpin id AA09310; Sun, 29 Sep 91 17:20:40 PDT
Received: by banshee.asd.sgi.com (910711.SGI/900721.SGI)
	for @giraffe.asd.sgi.com:mccalpin@perelandra.cms.udel.edu id AA08137; Sun, 29 Sep 91 17:20:38 -0700
Date: Sun, 29 Sep 91 17:20:38 -0700
From: blythe@banshee.asd.sgi.com (David Blythe)
Message-Id: <9109300020.AA08137@banshee.asd.sgi.com>
To: mccalpin
Subject: Re: Attainable memory bandwidth
Newsgroups: comp.benchmarks,comp.arch
In-Reply-To: <MCCALPIN.91Sep27212326@pereland.cms.udel.edu>
Organization: Silicon Graphics, Inc., Mountain View, CA
Cc: 
Status: RO

In article <MCCALPIN.91Sep27212326@pereland.cms.udel.edu> you write:
>==========================================================
>Memory Transfer Rates in MB/s for Standard Fortran Kernels
>==========================================================
>John D. McCalpin			September 24, 1991
>mccalpin@perelandra.cms.udel.edu	DELOCN::MCCALPIN
>
>I am ***very*** interested in results from the following:
>	HP 7x0 machines
>	SGI 4D/35
>	SGI Indigo
>	Sun SPARC 2
>I would appreciate any data that people could send along!

since you ask, here are the numbers for Indigo and for a 4D/440.  I didn't
see any handy 35s around, maybe someone else will send them.

SGI 4D/440 \w 256Mb 64K instr cache + 64K data cache + 1Mb secondary cache

 Timing calibration ; t =    35.99999     clicks
      
 Assignment: Rate =    44.44445     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    29.09091     MB/s MFLOPS =    1.818182    
 Summing:    Rate =    47.05882     MB/s MFLOPS =    1.960784    
 SAXPYing:   Rate =    32.87671     MB/s MFLOPS =    2.739726    

SGI Indigo \w 16Mb 32K instr cache  32K data cache

 Timing calibration ; t =    23.99999     clicks
      
 Assignment: Rate =    66.66669     MB/s MFLOPS =   0.0000000E+00
 Scaling:    Rate =    27.58620     MB/s MFLOPS =    1.724138    
 Summing:    Rate =    57.14288     MB/s MFLOPS =    2.380953    
 SAXPYing:   Rate =    34.78259     MB/s MFLOPS =    2.898550    

	david blythe
	blythe@sgi.com



From alz@grumpy.cray.com  Mon Sep 30 08:50:03 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA18122; Mon, 30 Sep 91 08:50:03 EDT
Received: from dopey.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6j)
	id AA17526; Mon, 30 Sep 91 07:48:27 CDT
Received: by dopey.cray.com
	id AA00630; 4.1/CRI-4.4; Mon, 30 Sep 91 08:48:25 EDT
Date: Mon, 30 Sep 91 08:48:25 EDT
From: alz@grumpy.cray.com (Andrew Zachary)
Message-Id: <9109301248.AA00630@dopey.cray.com>
To: mccalpin
Subject: Re:  Cray Results for bandwidth test
In-Reply-To: Mail from 'mccalpin@perelandra.cms.udel.edu'
      dated: Mon, 30 Sep 91 08:41:40 EDT
Cc: alz@grumpy.cray.com
Status: RO

John,

Your guess is correct: the optimizer is removing all those dead blocks
of code!  I put each of the loops into a subroutine, turned off double
precision, and then used second to time the loops.  The results now
make sense: 150 Mflops for the sum and scale, 300 Mflops for the saxpy,
and storage rates of 2400 Mbytes/sec for sum and scale and 3600 Mbytes/sec
on saxpy.  Note that the 3600 Mbytes/sec figure is a bit below the
theoretical peak of 4000 Mbytes/sec for a saxpy, but I wasn't running
on a dedicated machine.  Also, I had to run on a machine with only
64 Memory banks, so I may have had some unexpected bank conflicts.

I will send you a "certified" copy of the results once my machine
comes back up...

Andrew Zachary

P.S.  If you would like to see the results for a dedicated 8 processor
machine, let me know.  I can probably arrange the time.


From UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper  Mon Sep 30 11:47:56 1991
Received: from ucsd.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA18754; Mon, 30 Sep 91 11:47:56 EDT
Received: from celit.UUCP by ucsd.edu; id AA00664
	sendmail 5.64/UCSD-2.1-sun via UUCP
	Mon, 30 Sep 91 08:41:39 -0700
Received: by celit.fps.com (5.51/celerity1.1)
	id AA20241; Mon, 30 Sep 91 08:34:05 PDT for mccalpin@perelandra.cms.udel.edu at ucsd
Posted-Date: Mon, 30 Sep 91 08:32:33 PDT
Received: from keeper.fps_net by fpssun (4.1/SMI-4.1)
	id AA22799; Mon, 30 Sep 91 08:26:10 PDT
Received: by keeper.fps_net (4.1/SMI-4.1)
	id AA04146; Mon, 30 Sep 91 08:32:33 PDT
Date: Mon, 30 Sep 91 08:32:33 PDT
From: UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper (Brian Whitney)
Message-Id: <9109301532.AA04146@keeper.fps_net>
To: mccalpin
Subject: memory xfer rates program
Status: RO


Hi,

Brian Whitney with FPS Computing.

I tried running your memory xfer program on our Model 500
SPARC series computer and came up with results that
are a little concerning to me.

First of all, the type in your program is

>>>      parameter (N= 1 000 000)
>>>      real a(N),b(N),c(N)
>>>      real second

While in your test loop, you have a double precision constant.
This will cause extra operations in FORTRAN, consisting of type
conversions.

>>>     do 30 j=1,N
>>>        c(j) = 3.0d0*a(j)
>>>  30 continue

I believe you meant to have code run in 64-bit precision, because
results I achieved would say that is the case.

When I changed the Real declartion to real*8, the results look
correct to me for our machine.


Here are results for the FPS S511, SPARC based cpu with added vector
processor.

Code as presented...

systst2% x.x.
Timing calibration ; t =     6.45360 clicks
    
Assignment: Rate =     247.924 MB/s MFLOPS =   0.
Scaling:    Rate =     102.594 MB/s MFLOPS =     6.41215
Summing:    Rate =     186.564 MB/s MFLOPS =     7.77351
SAXPYing:   Rate =     127.367 MB/s MFLOPS =     10.6139


Code modified to make 3.0d0 -> 3.0 (pure 32 bit run, I believe the
Mbyte calculation is 2x what it should be...)

systst2% x.x
Timing calibration ; t =     6.53050 clicks
    
Assignment: Rate =     245.004 MB/s MFLOPS =   0.
Scaling:    Rate =     169.176 MB/s MFLOPS =     10.5735
Summing:    Rate =     186.296 MB/s MFLOPS =     7.76235
SAXPYing:   Rate =     185.753 MB/s MFLOPS =     15.4794


Code modified to have real -> real*8 (mbytes should be correct)

systst2% x8.x
Timing calibration ; t =     6.42171 clicks
    
Assignment: Rate =     249.155 MB/s MFLOPS =   0.
Scaling:    Rate =     168.101 MB/s MFLOPS =     10.5063
Summing:    Rate =     188.048 MB/s MFLOPS =     7.83534
SAXPYing:   Rate =     189.218 MB/s MFLOPS =     15.7681


The last data are math library routines replacing your loops.
This just shows closer to what the limits of the machine are...

systst2% x2.x
Timing calibration ; t =     6.45340 clicks
    
Assignment: Rate =     247.931 MB/s MFLOPS =   0.
Scaling:    Rate =     245.912 MB/s MFLOPS =     15.3695
Summing:    Rate =     241.687 MB/s MFLOPS =     10.0703
SAXPYing:   Rate =     244.302 MB/s MFLOPS =     20.3585


Brian Whitney
FPS Computing
3601 SW Murray Blvd
Beaverton, OR  97005

(503) 641-3151 x2682
keeper@fps.com

From brooks@tazdevil.llnl.gov  Mon Sep 30 12:54:03 1991
Received: from tazdevil.llnl.gov by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA18899; Mon, 30 Sep 91 12:54:03 EDT
Received: by tazdevil.llnl.gov (4.1/1.15)
	id AA28890; Mon, 30 Sep 91 09:52:41 PDT
Date: Mon, 30 Sep 91 09:52:41 PDT
From: brooks@tazdevil.llnl.gov (Eugene D. Brooks III)
Message-Id: <9109301652.AA28890@tazdevil.llnl.gov>
To: mccalpin
Subject: Re:  your kernels on a SPARC-2
Status: RO

The code was compiled with -O on a SOLBOURNE.

From lucier@newton.math.purdue.edu  Mon Sep 30 14:54:50 1991
Received: from gauss.math.purdue.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA19659; Mon, 30 Sep 91 14:54:50 EDT
Received: from newton.math.purdue.edu.xanth by gauss.math.purdue.edu (4.1/1.13jrs)
	id AA04086; Mon, 30 Sep 91 13:53:11 EST
Received: by newton.math.purdue.edu.xanth (4.1/SMI-4.1)
	id AA28927; Mon, 30 Sep 91 13:52:54 EST
Date: Mon, 30 Sep 91 13:52:54 EST
From: lucier@newton.math.purdue.edu (Bradley Lucier)
Message-Id: <9109301852.AA28927@newton.math.purdue.edu.xanth>
To: mccalpin
Subject: stream_s on HP
Cc: lucier@newton.math.purdue.edu
Status: RO

I wouldn't say I'm a fan of the HP 720 (not yet, at least),
but in compiling your code there was not a function etime to
get the time.  (The loader complained.)
I don't have a complete set of man pages, etc.,
with our evaluation version of the machine that goes away Wednesday
morning, and I'm not really familiar with SYS V, so I don't know
what routine to substitute on the HP.  Would you be willing to
hazard a guess?

Brad Lucier lucier@math.purdue.edu

From lucier@newton.math.purdue.edu  Mon Sep 30 15:24:27 1991
Received: from gauss.math.purdue.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA19763; Mon, 30 Sep 91 15:24:27 EDT
Received: from newton.math.purdue.edu.xanth by gauss.math.purdue.edu (4.1/1.13jrs)
	id AA04340; Mon, 30 Sep 91 14:22:49 EST
Received: by newton.math.purdue.edu.xanth (4.1/SMI-4.1)
	id AA28964; Mon, 30 Sep 91 14:22:32 EST
Date: Mon, 30 Sep 91 14:22:32 EST
From: lucier@newton.math.purdue.edu (Bradley Lucier)
Message-Id: <9109301922.AA28964@newton.math.purdue.edu.xanth>
To: mccalpin
Subject: Re:  stream_s on HP
Cc: lucier@newton.math.purdue.edu
Status: RO

Here's what I got for the single precision version.  Something is screwed
up for the double.

[6] % f77 -O -o stream_s stream_s.f second.o
stream_s.f:
   MAIN stream:
   realsize:
   dummy:
[7] % stream_s
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time =  35.0 hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   33.3334       .1261       .1200       .1300  
Scaling   :   30.7693       .1371       .1300       .1400  
Summing   :   35.2942       .1741       .1700       .1800  
SAXPYing  :   31.5790       .1920       .1900       .2000  


I can't tell you what the OS version, compiler version, etc, are.
Here are the double precision results.  Any suggestions?

[10] % !!
stream_d 
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =  5.981443217023579E-03 hundredths
  of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:13700.9031   9772.4978       .0004  23593.0000  
Scaling   : 1471.6167  11931.8747       .0033  26214.3750  
Summing   :  409.6001  14189.2160       .0176  30146.5625  
SAXPYing  :  112.9412  20303.4971       .0637  52428.7500  

From lucier@newton.math.purdue.edu  Mon Sep 30 16:59:27 1991
Received: from gauss.math.purdue.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA20507; Mon, 30 Sep 91 16:59:27 EDT
Received: from newton.math.purdue.edu.xanth by gauss.math.purdue.edu (4.1/1.13jrs)
	id AA05122; Mon, 30 Sep 91 15:57:51 EST
Received: by newton.math.purdue.edu.xanth (4.1/SMI-4.1)
	id AA28977; Mon, 30 Sep 91 15:57:34 EST
Date: Mon, 30 Sep 91 15:57:34 EST
From: lucier@newton.math.purdue.edu (Bradley Lucier)
Message-Id: <9109302057.AA28977@newton.math.purdue.edu.xanth>
To: mccalpin
Subject: Re:  stream_s on HP
Cc: lucier@newton.math.purdue.edu
Status: RO

OK, with the same caveats, here are the double precision results:

[6] % f77 -O -o stream_d stream_d.f second.o
stream_d.f:
   MAIN stream:
   realsize:
   dummy:
[7] % stream_d
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =  35.99999919533729 hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   43.6365       .1161       .1100       .1200  
Scaling   :   40.0000       .1200       .1200       .1200  
Summing   :   45.0000       .1600       .1600       .1600  
SAXPYing  :   44.9999       .1690       .1600       .1700  
[8] % 

From abe@mace.cc.purdue.edu  Mon Sep 30 17:23:17 1991
Received: from mace.cc.purdue.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA20581; Mon, 30 Sep 91 17:23:17 EDT
Received: from localhost by mace.cc.purdue.edu (5.61/Purdue_CC)
	id AA17317; Mon, 30 Sep 91 16:21:46 -0500
Message-Id: <9109302121.AA17317@mace.cc.purdue.edu>
To: mccalpin
Subject: stream_[ds] results for HP 9000/720
Date: Mon, 30 Sep 91 16:21:45 -0500
From: abe@mace.cc.purdue.edu
Status: RO

I am not an HP-UX expert -- I just happen to have access to a 9000/720
at the moment --  so I'm not sure I exploited the compiler to its fullest.
I did try some exotic optimization flags and couldn't get the results to
execute at all.

There were no man pages for any timing functions on this system (including
etime(), dtime(), mtime() or even times()).  I had to use FORTRAN/C functions
that I concocted for Linpack to come up with a second().  They usually
work everywhere, because they start with a FORTRAN second() function that
calls a C function, capable of using times() or getrusage().

Vic Abell <abe@mace.cc.purdue.edu>

====================

Hardware: HP 9000/720

OS: HP-UX A.B8.05 (as reported by ``uname -r'')

Compiler: HP FORTRAN 77, Ver: 8.05

Compiler options: f77 +O3

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =  37.0 hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   43.6364       .1171       .1100       .1200  
Scaling   :   40.0000       .1221       .1200       .1300  
Summing   :   45.0000       .1620       .1600       .1700  
SAXPYing  :   42.3529       .1720       .1700       .1800  
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time =  35.0 hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   33.3334       .1281       .1200       .1300  
Scaling   :   30.7693       .1341       .1300       .1400  
Summing   :   35.2942       .1761       .1700       .1800  
SAXPYing  :   31.5790       .1900       .1900       .1900  

From uunet.UU.NET!stardent!Stardent.COM!wright  Mon Sep 30 17:34:47 1991
Received: from relay2.UU.NET by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA20589; Mon, 30 Sep 91 17:34:47 EDT
Received: from uunet.uu.net (via LOCALHOST.UU.NET) by relay2.UU.NET with SMTP 
	(5.61/UUNET-internet-primary) id AA27897; Mon, 30 Sep 91 17:33:22 -0400
Received: from stardent.UUCP by uunet.uu.net with UUCP/RMAIL
	(queueing-rmail) id 173239.13091; Mon, 30 Sep 1991 17:32:39 EDT
Received: by stardent.Stardent.COM (1.1/smail2.5/01-28-89)
	id AA04921; Mon, 30 Sep 91 17:24:32 EDT
Message-Id: <9109302124.AA04921@stardent.Stardent.COM>
To: uunet.UU.NET!uunet!perelandra.cms.udel.edu!mccalpin
Subject: Re: Attainable memory bandwidth
Date: Mon, 30 Sep 91 17:24:28 EDT
From: David Wright <wright@Stardent.COM>
Status: RO

You write:

>I inadvertantly transmitted a bad copy of the program with the
>referenced message.  Due to several errors, this code gives bogus
>results, and should be discarded.  The results in the table that I
>posted are not too bad, but the results from running the program that
>I actually sent out are worthless.

>I have built two new versions that (I hope) are much more reliable.
>They are called stream_s.f and stream_d.f and are available by
>anonymous ftp from perelandra.cms.udel.edu in bench/stream/.

Gee, and here we'd already been playing around with it and discovering
that it was bad.  We were all set to email you...

Anyway, Stardent doesn't have ftp capability, but we'd love have
copies of your (corrected) programs, particularly to see how our 
Vistra machines do.  So, could I prevail upon you to email copies?
We'll email the results back when we get 'em.  After all, you really
should have an i860-based machine in your results.

  -- David Wright, Stardent Computer Inc
     wright@stardent.com

From olson@anchor.esd.sgi.com  Mon Sep 30 19:49:11 1991
Received: from SGI.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA20661; Mon, 30 Sep 91 19:49:11 EDT
Received: from [192.26.58.38] by sgi.sgi.com via SMTP (910911.SGI.EXPERIMENTAL/910110.SGI)
	for mccalpin@perelandra.cms.udel.edu id AA02583; Mon, 30 Sep 91 16:47:35 -0700
Received: by anchor.esd.sgi.com (910711.SGI/910805.SGI)
	for @sgi.sgi.com:mccalpin@perelandra.cms.udel.edu id AA19208; Mon, 30 Sep 91 16:47:34 -0700
Date: Mon, 30 Sep 91 16:47:34 -0700
From: olson@anchor.esd.sgi.com (Dave Olson)
Message-Id: <9109302347.AA19208@anchor.esd.sgi.com>
To: "John D. McCalpin" <mccalpin>
Subject: re: Memory bandwidth testing
Status: RO

Here are the results from the 35 and the indigo, both running 4.0
with 16 Mbytes of memory, compiled as f77 -O3.  I ran it 3 times
in a row on each.

=========== out.SGI_35_s =====
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time =    50.00000     hundredths of a second
 Timing calibration ; time =    53.00000     hundredths of a second
 Timing calibration ; time =    52.00000     hundredths of a second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   50.0000      0.0875      0.0800      0.1100
          :   50.0000      0.0832      0.0800      0.1000
          :   50.0000      0.0831      0.0800      0.0900
Scaling   :   28.5715      0.1452      0.1400      0.1600
          :   28.5715      0.1421      0.1400      0.1500
          :   28.5715      0.1481      0.1400      0.1600
Summing   :   37.5000      0.1661      0.1600      0.1800
          :   37.5001      0.1620      0.1600      0.1700
          :   37.5000      0.1661      0.1600      0.1800
SAXPYing  :   25.0000      0.2538      0.2400      0.3100
          :   25.0000      0.2410      0.2400      0.2500
          :   25.0000      0.2501      0.2400      0.2600
=========== out.SGI_35_d =====
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    45.99999990314245     hundredths of a second
 Timing calibration ; time =    45.00000104308128     hundredths of a second
 Timing calibration ; time =    45.00000085681677     hundredths of a second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   53.3334      0.1129      0.0900      0.2000
          :   53.3335      0.0910      0.0900      0.1000
          :   53.3335      0.0900      0.0900      0.0900
Scaling   :   36.9233      0.1352      0.1300      0.1500
          :   36.9231      0.1351      0.1300      0.1400
          :   36.9231      0.1381      0.1300      0.1500
Summing   :   48.0001      0.1698      0.1500      0.2000
          :   45.0000      0.1632      0.1600      0.1900
          :   45.0000      0.1641      0.1600      0.1700
SAXPYing  :   42.3530      0.1834      0.1700      0.2100
          :   42.3530      0.1825      0.1700      0.2200
          :   42.3530      0.1842      0.1700      0.2000
=========== out.SGI_Indigo_s ====
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time =    58.00000     hundredths of a second
 Timing calibration ; time =    60.00000     hundredths of a second
 Timing calibration ; time =    61.00000     hundredths of a second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.5714      0.1512      0.1400      0.1600
          :   28.5715      0.1462      0.1400      0.1600
          :   28.5715      0.1463      0.1400      0.1700
Scaling   :   20.0000      0.2184      0.2000      0.2400
          :   21.0526      0.2136      0.1900      0.2400
          :   21.0526      0.2114      0.1900      0.2300
Summing   :   25.0000      0.2574      0.2400      0.2800
          :   26.0870      0.2463      0.2300      0.2700
          :   26.0869      0.2525      0.2300      0.2800
SAXPYing  :   18.1818      0.3637      0.3300      0.3900
          :   18.7500      0.3536      0.3200      0.3800
          :   18.1818      0.3523      0.3300      0.3800
========== out.SGI_Indigo_d ====
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    53.00000030547380     hundredths of a second
 Timing calibration ; time =    50.99999718368053     hundredths of a second
 Timing calibration ; time =    56.00000526756048     hundredths of a second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   32.0000      0.1683      0.1500      0.2200
          :   34.2859      0.1647      0.1400      0.1900
          :   32.0000      0.1656      0.1500      0.2000
Scaling   :   22.8571      0.2378      0.2100      0.2800
          :   22.8571      0.2346      0.2100      0.2700
          :   22.8571      0.2396      0.2100      0.2700
Summing   :   27.6924      0.2936      0.2600      0.3500
          :   27.6923      0.2849      0.2600      0.3300
          :   28.8000      0.2785      0.2500      0.3100
SAXPYing  :   26.6667      0.3269      0.2700      0.4000
          :   25.7143      0.3079      0.2800      0.3600
          :   25.7143      0.3015      0.2800      0.3400

     Dave Olson
----============----
Life would be so much easier if we could just look at the source code.


From hudgens@sun13.scri.fsu.edu  Mon Sep 30 20:48:22 1991
Received: from sun13.scri.fsu.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA20684; Mon, 30 Sep 91 20:48:22 EDT
Received:  by sun13.scri.fsu.edu (5.61/25-eef)
	id AA25290; Mon, 30 Sep 91 20:44:50 -0400
Date: Mon, 30 Sep 91 20:44:50 -0400
From: Jim Hudgens <hudgens@sun13.scri.fsu.edu>
Message-Id: <9110010044.AA25290@sun13.scri.fsu.edu>
To: mccalpin (John D. McCalpin)
In-Reply-To: mccalpin@perelandra.cms.udel.edu's message of 30 Sep 91 18:38:08 GMT
Subject: Re: Attainable memory bandwidth
Status: RO



  By the way, I have not heard from anyone with an HP-720 or HP-730.
  Don't all you HP fans out there want to show off?

Login to demo453a.scri.fsu.edu.  I believe you have an account there.
It's a 720 with 32M of memory.

JHH
Jim Hudgens			Supercomputer Computations Research Institute
"Nothing's for sure except DEC and VAXe .. er, death and taxes"
hudgens@sun13.scri.fsu.edu	Life's a bitch, and then you graduate.

From csrcb@shark.mel.dit.csiro.au  Mon Sep 30 21:48:09 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA20708; Mon, 30 Sep 91 21:48:09 EDT
Received: from trout.mel.dit.CSIRO.AU by shark.mel.dit.csiro.au with SMTP id AA03114
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Tue, 1 Oct 91 11:46:34 +1000
Received: by trout.mel.dit.CSIRO.AU (4.1/SMI-4.0)
	id AA12175; Tue, 1 Oct 91 11:46:33 EST
Date: Tue, 1 Oct 91 11:46:33 EST
From: Robert.Bell@mel.dit.csiro.au
Message-Id: <9110010146.AA12175@trout.mel.dit.CSIRO.AU>
To: mccalpin
Subject: stream_s.f
Status: RO

 John,
      I ran stream_s.f on our Cray - the compiler optimized out all
 the interesting code again.
      Rob.
 P.S.  The simple insertion of
      Common // a, b, c
 enabled correct results to be achieved.  I also used the Cray intrinsic
 second, which returns CPU time - appropriate on a busy machine.
 Here are the results.

 Run on Cray Y-MP2/216, serial number 1409.
 Clock speed 5.998 ns.
 Banks are busy for 5 clock periods per reference.
 Using only one processor.
 UNICOS 6.0.12
 Compiler invocation: cf77 -V -o Mc -Wf"-a static" stream_s.f  >& stream_s.out
 Executable invocation: ./Mc >> stream_s.out
 1991 Oct 01  11:36:55 Tue

/ Robert C. Bell		     |  CSIRO Supercomputing Support Manager  \
| Division of Information Technology | 'phone: (03) 282 2620   +61 3 282 2620 |
| 723 Swanston Street		     |	fax:   (03) 282 2600   +61 3 282 2600 |
\ Carlton VIC 3053 Australia	     |	email: csrcb@mel.dit.csiro.au	      /


 FF0001 CFT77 VERSION 4.0.3    (386394) 12/22/90 21:03:26                     
 FF0002 COMPILE TIME     .563 SECONDS                                         
 FF0006 MAXIMUM FIELD LENGTH   333805 DECIMAL WORDS                           
 FF0003 266 SOURCE LINES                                                      
 FF0004 0 ERRORS, 0 WARNINGS                                                  
 FF0005 CODE: 378 WORDS, DATA: 392 WORDS                                      
 SEGLDR version 6.0 - 08/25/91  (Chg-01/21/91  Level-0)
 (c) Copyright Cray Research, Inc.
 Unpublished -- All rights reserved under copyright laws of the United States
--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 0.9858252 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment: 2426.3556      0.0033      0.0033      0.0033  
Scaling   : 2426.1790      0.0033      0.0033      0.0033  
Summing   : 3454.4045      0.0035      0.0035      0.0036  
SAXPYing  : 3396.8949      0.0036      0.0035      0.0036  

From csrcb@shark.mel.dit.csiro.au  Mon Sep 30 20:14:26 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA20677; Mon, 30 Sep 91 20:14:26 EDT
Received: from trout.mel.dit.CSIRO.AU by shark.mel.dit.csiro.au with SMTP id AA00863
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Tue, 1 Oct 91 10:12:48 +1000
Received: by trout.mel.dit.CSIRO.AU (4.1/SMI-4.0)
	id AA11499; Tue, 1 Oct 91 00:12:08 EST
Date: Tue, 1 Oct 91 00:12:08 EST
From: Robert.Bell@mel.dit.csiro.au
Message-Id: <9109301412.AA11499@trout.mel.dit.CSIRO.AU>
To: mccalpin
Status: RO

 John,
      I tried your benchmark on a Cray Y-MP2/216, using just a single processor.
 The cycle time is 6.0 ns (near enough).
      I needed to make several modifications.  I firstly deleted your second
 function, and used Cray's supplied intrinsic.  I then executed the program,
 and obtained sensational but erroneous results.  Because your program does not
 use the array c, although defining it several times, the compiler seemed to
 optimize out completely the do loops - not a bad feature.  I inserted some
 code to force the storage of the arrays, and turned on a compiler feature
 to convert double precision features to real (64-bit of course on the Cray),
 and eventually received the following results.

 Timing calibration ; t = 0.6722976 clicks
     
 Assignment: Rate = 2379.898425935 MB/s MFLOPS = 0.
 Scaling:    Rate = 2335.76778731 MB/s MFLOPS = 145.9854867069
 Summing:    Rate = 3220.624881743 MB/s MFLOPS = 134.1927034059
 SAXPYing:   Rate = 3264.136362561 MB/s MFLOPS = 272.0113635467

    The theoretical peaks are 2666 Mbyte/s for the first two, and
 4000 Mbyte/s for the second two tests.  (2*8/6.0e-9 and 3*8/6.0e-9) / 1.0e6.

    I have been interested in memory transfer times for several years now.
 I have run similar tests to yours, but also looking at relative memory
 locations for the arrays, and scatter-gathers with various strides.   On some
 machines, the relative locations of the arrays (whether they start in the same
 bank or not, etc) can have a substantial effect.
     Incidentally, I have some results from the Cray C-90 (=Y-MP16).  The best 
 result for a loop like your assignment loop is 6937 Mbyte /s.
 
     Here are some speeds for various machines, expressed in Mbyte/s.  The first
 group is for a(i) = b(i+k-1), k = 1, 128, and a and b being forced to start
 in the same bank.  The two figures for each machine represent the slowest and
 fastest times seen as k varied.
     The second group is for loops like a(i) = b(1 + k*(i-1)) for various values
 of k.
     I will be happy to provide more information - I would like to eventually
 publish the improved version of my benchmark code, which I am working on.


'a(i) = b(i+k-1) - offset$'
'C-90$', 6511., 6937. /
'Y-MP$', 2340., 2423. /
'X-MP$',  891., 1686. /
'205$',  1447., 1587. /
'C220$',  137.,  150. /
'ETA 10P*$', 1400., 1506. /
'ETA 10P$',  619., 1305. /
'SG 280$',  17.7, 19.3 /
'$' / 
'scatter-gather - stride$'
'C-90$', 533., 6936. /
'Y-MP$', 511., 2412. /
'X-MP$',  232., 1506. /
'205$',  185., 699. /
'C220$',  31.,  162. /
'ETA 10P*$', 73., 565. /
'ETA 10P$',  25., 474. /
'SG 280$',  5.2, 16.4 /
'$' / 

 Robert Bell.

/ Robert C. Bell		     |  CSIRO Supercomputing Support Manager  \
| Division of Information Technology | 'phone: (03) 282 2620   +61 3 282 2620 |
| 723 Swanston Street		     |	fax:   (03) 282 2600   +61 3 282 2600 |
\ Carlton VIC 3053 Australia	     |	email: csrcb@mel.dit.csiro.au	      /

From Daniel.Stodolsky@ERNST.MACH.CS.CMU.EDU  Tue Oct  1 09:26:16 1991
Received: from ERNST.MACH.CS.CMU.EDU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA21575; Tue, 1 Oct 91 09:26:16 EDT
Received: from Messages.7.8.N.CUILIB.3.45.SNAP.NOT.LINKED.ERNST.MACH.CS.CMU.EDU.pmax.3
          via MS.5.6.ERNST.MACH.CS.CMU.EDU.pmax_3;
          Tue,  1 Oct 91 09:24:12 -0400 (EDT)
Message-Id: <Ecu7Pwy00h70MByrQ4@cs.cmu.edu>
Date: Tue,  1 Oct 91 09:24:12 -0400 (EDT)
From: Daniel.Stodolsky@cs.cmu.edu
To: "John D. McCalpin" <mccalpin>
Subject: Re: Attainable memory bandwidth
In-Reply-To: <MCCALPIN.91Sep30143808@pereland.cms.udel.edu>
References: <MCCALPIN.91Sep27212326@pereland.cms.udel.edu>,
	<MCCALPIN.91Sep30143808@pereland.cms.udel.edu>
Status: RO

This is for an 
Omron Luna88k SX-9100/DT quad processor, 40 Mbytes of memory. The
compiler does not parallelize code. 
Running Mach 2.5.

Danner


Script started on Tue Oct  1 09:19:31 1991
amalia[1] f77 -v -OLM stream_d.f
Green Hills Software, Inc.  Fortran Compiler  Version 1.8.5 November 1,
1990
/usr/lib/ccom/fcom88 -X85 -X71 stream_d.f stream_d.s -OLM
/bin/as stream_d.s -o stream_d.o
/bin/ld -X /usr/lib/crt0.o stream_d.o -lF77 -lI77 -lU77 -lm -ltermcap
-lindep -lc
rm stream_d.s stream_d.o
amalia[2] ./a.out
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
  Timing calibration ; time =    149.99999329448   hundredths   of a
second
  Increase the size of the arrays if this is <30 
   and your clock precision is =<1/100 second
  ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   14.4000      0.3384      0.3333      0.3500
Scaling   :   14.4000      0.3521      0.3333      0.4000
Summing   :   16.0000      0.4567      0.4500      0.4667
SAXPYing  :   13.0909      0.5551      0.5500      0.5667
amalia[3] f77 -v -OLM stream_s.f
Green Hills Software, Inc.  Fortran Compiler  Version 1.8.5 November 1,
1990
/usr/lib/ccom/fcom88 -X85 -X71 stream_s.f stream_s.s -OLM
/bin/as stream_s.s -o stream_s.o
/bin/ld -X /usr/lib/crt0.o stream_s.o -lF77 -lI77 -lU77 -lm -ltermcap
-lindep -lc
rm stream_s.s stream_s.o
amalia[4] ./a.out
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
  Timing calibration ; time =    121.667   hundredths   of a second
  Increase the size of the arrays if this is <30 
   and your clock precision is =<1/100 second
  ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   14.1177      0.2919      0.2833      0.3167
Scaling   :   13.3334      0.3000      0.3000      0.3000
Summing   :   15.6522      0.3901      0.3833      0.4000
SAXPYing  :   13.3333      0.4601      0.4500      0.4667
amalia[5] exit
amalia[6] 
script done on Tue Oct  1 09:20:58 1991


From @scapa.cs.ualberta.ca:kufeld_k@edhp01  Tue Oct  1 15:06:53 1991
Received: from [129.128.4.44] by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA01220; Tue, 1 Oct 91 15:06:53 EDT
Received: from ipl by scapa.cs.ualberta.ca with UUCP id <42591>; Tue, 1 Oct 1991 13:05:08 -0600
Received: by edhp01
	(16.7/16.2) id AA01034; Tue, 1 Oct 91 12:30:59 -0600
From: Kurt Kufeld <kufeld_k@edhp01.uucp>
Subject: Memory bandwidth benchmarks
To: mccalpin
Date: 	Tue, 1 Oct 1991 12:30:58 -0600
Mailer: Elm [revision: 66.33]
Message-Id: <91Oct1.130508mdt.42591@scapa.cs.ualberta.ca>
Status: RO

John:

	I've got an HP9000/730 and am willing to show off.  I do not have
a direct connect to the Internet and would require your source for the
benchmark mailed to me.  If you want me to test, then mail it.

				Kurt Kufeld
				kufeld_k%ipl@cs.ualberta.ca
				(403) 420-5330

From UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper  Tue Oct  1 16:41:16 1991
Received: from ucsd.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA01739; Tue, 1 Oct 91 16:41:16 EDT
Received: from celit.UUCP by ucsd.edu; id AA20920
	sendmail 5.64/UCSD-2.2-sun via UUCP
	Tue, 1 Oct 91 13:37:47 -0700
Received: by celit.fps.com (5.51/celerity1.1)
	id AA14766; Tue, 1 Oct 91 13:35:28 PDT for mccalpin@perelandra.cms.udel.edu at ucsd
Posted-Date: Tue, 1 Oct 91 13:30:03 PDT
Received: from keeper.fps_net by fpssun (4.1/SMI-4.1)
	id AA05271; Tue, 1 Oct 91 13:23:39 PDT
Received: by keeper.fps_net (4.1/SMI-4.1)
	id AA05751; Tue, 1 Oct 91 13:30:03 PDT
Date: Tue, 1 Oct 91 13:30:03 PDT
From: UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper (Brian Whitney)
Message-Id: <9110012030.AA05751@keeper.fps_net>
To: mccalpin
Subject: re: Memory bandwidth testing
Status: RO


I would very much like to run your benchmark on the platforms I
have available to me.  Unfortunately, my FTP access is messed up
at the moment.

Please mail me the relevant files and I will run them on the
machines I have access to.

Brian Whitney
FPS Computing 

keeper@fps.com

From csrcb@shark.mel.dit.csiro.au  Tue Oct  1 19:03:00 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02015; Tue, 1 Oct 91 19:03:00 EDT
Received: from trout.mel.dit.CSIRO.AU by shark.mel.dit.csiro.au with SMTP id AA02628
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Tue, 1 Oct 91 10:47:03 +1000
Received: by trout.mel.dit.CSIRO.AU (4.1/SMI-4.0)
	id AA12020; Tue, 1 Oct 91 10:47:02 EST
Date: Tue, 1 Oct 91 10:47:02 EST
From: Robert.Bell@mel.dit.csiro.au
Message-Id: <9110010047.AA12020@trout.mel.dit.CSIRO.AU>
To: mccalpin
Subject: Benchmark
Status: RO

 John,
      Because I have nothing netter to do (ha ha), I have fiddled with your
 benchmark further, to try to build defences against aggressive optimization.
      You are welcome to incorporate any of my changes, with acknowledgement.
      Regards,
              Rob. Bell.
/ Robert C. Bell		     |  CSIRO Supercomputing Support Manager  \
| Division of Information Technology | 'phone: (03) 282 2620   +61 3 282 2620 |
| 723 Swanston Street		     |	fax:   (03) 282 2600   +61 3 282 2600 |
\ Carlton VIC 3053 Australia	     |	email: csrcb@mel.dit.csiro.au	      /


        program Mc2
C
C       This program tests memory transfers.
C       Original by
C John D. McCalpin                        mccalpin@perelandra.cms.udel.edu
C Assistant Professor                     mccalpin@brahms.udel.edu
C College of Marine Studies, U. Del.      DELOCN::MCCALPIN (SPAN)
C
C       Modified by
C Robert C. Bell		     |  CSIRO Supercomputing Support Manager  \
C Division of Information Technology | 'phone: (03) 282 2620   +61 3 282 2620 |
C 723 Swanston Street		     |	fax:   (03) 282 2600   +61 3 282 2600 |
C Carlton VIC 3053 Australia	     |	email: csrcb@mel.dit.csiro.au	      /
C
C       1991 Oct 01  09:47:24 Tue
C
C Warning - machine dependencies.
C To run correctly, reals must be 64-bit, integers must be able to hold
C numbers up to a million, and a function called second must be available.
C The real function second, which has no arguments, returns the CPU seconds
C used so far in the program.
C This is preferable to using elapsed time for busy machines.
C
	parameter (N= 1 000 000)
	real a(N),b(N),c(N)
	real second, t1, t2, t3, t4, toverh, t
	real check, const
C
C       Put the arrays in common, to force the same alignment.
        common // a, b, c
C
        data const / 3.0 /
        data check / 0.0 /

C
C       Initialize the arrays with non-trivial values, to stop
C       the simple optimization of just doing the calculation for
C       one array value.
C       Assume that integers as large as one million can be handled.
	do 10 j=1,N
	    a(j) = j
	    b(j) = N  -j
	    c(j) = - j
   10	continue
        call sub (N, a, b, c, check)
C
C       Insert a few extra calls to second, to by-pass initialization
C       which may distort the first few calls.
        t1 = second ( )
        t2 = second ( )
        t3 = second ( )
        t4 = second ( )
        toverh = t4 - t3
        print *, ' The first few times are ', t1, t2, t3, t4
        print *, ' The overhead time is about ', toverh


	t = second( )
	do 20 j=1,N
	    c(j) = a(j)
   20	continue
	t = second( )-t - toverh
        call sub (N, a, b, c, check)
	print *,'Timing calibration ; t = ',t*100,' clicks'
	print *,'    '
	print *,'Assignment: Rate = ',1.0e-6*real(2*N*8)/t,' MB/s'
     $         ,' MFLOPS = ',1.0e-6* real (0*N)/t

	t = second( )
	do 30 j=1,N
	    c(j) = const*a(j)
   30	continue
	t = second( )-t - toverh
        call sub (N, a, b, c, check)
	print *,'Scaling:    Rate = ',1.0e-6*real(2*N*8)/t,' MB/s'
     $         ,' MFLOPS = ',1.0e-6 * real (1*N)/t

	t = second( )
	do 40 j=1,N
	    c(j) = a(j)+b(j)
   40	continue
	t = second( )-t - toverh
        call sub (N, a, b, c, check)
	print *,'Summing:    Rate = ',1.0e-6*real(3*N*8)/t,' MB/s'
     $         ,' MFLOPS = ',1.0e-6 * real (1*N)/t

	t = second( )
	do 50 j=1,N
	    c(j) = a(j)+const*b(j)
   50	continue
	t = second( )-t - toverh
        call sub (N, a, b, c, check)
	print *,'SAXPYing:   Rate = ',1.0e-6*real(3*N*8)/t,' MB/s'
     $         ,' MFLOPS = ',1.0e-6 * real (2*N)/t
        print *, ' The check sum is ', check, 
     1    ', and should be 8999998.'
	end
        subroutine sub (N, a, b, c, check)
C
C       This routine forces the storage in the arrays.
C
C       It uses the value of pi to pseudo-randomly pick values
C       from the array to be used.
C       This may defeat most optimization attempts which try to not
C       compute all the array values.
C
        Integer N
        real a(n), b(n), c(n)
        real check
C
        real pi, rand
        integer icall, irand
        save pi, rand, icall, irand
C
        data icall / 0 /
C
        if (icall .le. 0) then
           pi = 4.0 * atan (1.0)
        endif
        icall = icall + 1
        rand = mod ( real (icall) * pi, 1.0)
        irand = int (rand * real (N))
        irand = min (N, max (1, irand))
        check = check + a(irand) + b(irand) + c(irand)
C No        write (*,*) pi, icall, rand, irand, 
C No     1     a(irand), b(irand), c(irand), check
        return
        end
C On Crays, using the provided intrinsic.
C No	real function second()
C No	real dummy(2)
C No	second = etime(dummy)
C No	end

From alz@grumpy.cray.com  Tue Oct  1 19:48:44 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02026; Tue, 1 Oct 91 19:48:44 EDT
Received: from grumpy.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6j)
	id AA21869; Tue, 1 Oct 91 18:47:08 CDT
Received: by grumpy.cray.com
	id AA18728; 4.1/CRI-4.4; Tue, 1 Oct 91 19:47:06 EDT
Date: Tue, 1 Oct 91 19:47:06 EDT
From: alz@grumpy.cray.com (Andrew Zachary)
Message-Id: <9110012347.AA18728@grumpy.cray.com>
To: mccalpin
Subject: Re:  Cray Results for bandwidth test
Cc: alz@grumpy.cray.com
Status: RO

John,

Even with your new program, I had to put each of your loops into
a separate subroutine to fool the compiler.  When I did that, I
ran the resulting code with each array = 2Mwords.  I ran on
sn1061, a Y-MP8/8128 with 6 nsec clock, 256 banks. With 1 cpu,
results are

--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 3.9378048 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment: 2436.8420      0.0131      0.0131      0.0131  
Scaling   : 2436.9222      0.0131      0.0131      0.0131  
Summing   : 3643.9862      0.0137      0.0132      0.0140  
SAXPYing  : 3517.9324      0.0138      0.0136      0.0140  


With 8 cpu's, the results are (cf77 -Zp)

--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 0.7495896003093 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:19291.5887      0.0019      0.0017      0.0029  
Scaling   :19294.1710      0.0019      0.0017      0.0028  
Summing   :26588.9383      0.0020      0.0018      0.0026  
SAXPYing  :26802.1963      0.0021      0.0018      0.0028  

Again, I am a little suprised that the I don't get 3 full words of I/O
per clock-cycle (read that as memory access).  I seem to get about 2.7
words/clock-period.  I will retry it using the BLAS level 1 routines, 
and see what I get.

Hope these results help you.
Andrew Zachary
alz@grumpy.cray.com

From Keith.Bierman@Eng.Sun.COM  Tue Oct  1 21:13:21 1991
Received: from sun.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02135; Tue, 1 Oct 91 21:13:21 EDT
Received: from Eng.Sun.COM (zigzag-bb.Corp.Sun.COM) by Sun.COM (4.1/SMI-4.1)
	id AA25918; Tue, 1 Oct 91 18:11:52 PDT
Received: from chiba.Eng.Sun.COM by Eng.Sun.COM (4.1/SMI-4.1)
	id AA20793; Tue, 1 Oct 91 18:11:44 PDT
Received: from localhost by chiba.Eng.Sun.COM (4.1/SMI-4.1)
	id AA02493; Tue, 1 Oct 91 18:11:41 PDT
Message-Id: <9110020111.AA02493@chiba.Eng.Sun.COM>
To: "John D. McCalpin" <mccalpin>
Subject: Re: Memory bandwidth testing 
In-Reply-To: Your message of Mon, 30 Sep 91 11:58:50 -0400.
             <9109301558.AA18806@perelandra.cms.udel.edu> 
Date: Tue, 01 Oct 91 18:11:38 PDT
From: Keith.Bierman@Eng.Sun.COM
Status: RO


Here is output from chiba, as before running 4.1.1, f77v1.4 patch
level 2, -fast -O4 -Bstatic, 5 runs. This time the window system is up
and running, some version of Openwindows rev3. 

Hopefully, one of my sunbuddies will be able to provide you with 4/6xx
series figures .... as those machines were announced yesterday.

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     99.999996274710 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2354      0.1791      0.1700      0.1900
Scaling   :   30.0000      0.1696      0.1600      0.2100
Summing   :   28.8000      0.2571      0.2500      0.2700
SAXPYing  :   26.6667      0.2791      0.2700      0.3000
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     97.999998182058 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   30.0000      0.1711      0.1600      0.1800
Scaling   :   30.0000      0.1661      0.1600      0.1800
Summing   :   28.8000      0.2530      0.2500      0.2600
SAXPYing  :   26.6667      0.2740      0.2700      0.2800
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     97.999998182058 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2353      0.1814      0.1700      0.2000
Scaling   :   30.0000      0.1762      0.1600      0.1900
Summing   :   28.8000      0.2694      0.2500      0.2900
SAXPYing  :   26.6667      0.2925      0.2700      0.3200
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =    104.999991506338 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2354      0.1843      0.1700      0.2000
Scaling   :   30.0000      0.1773      0.1600      0.2000
Summing   :   28.8000      0.2695      0.2500      0.2900
SAXPYing  :   26.6667      0.2885      0.2700      0.3200
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     107.00000151992 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2353      0.1824      0.1700      0.2000
Scaling   :   30.0000      0.1793      0.1600      0.2000
Summing   :   28.8000      0.2672      0.2500      0.2800
SAXPYing  :   26.6666      0.2883      0.2700      0.3100
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     95.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1819      0.2251      0.2200      0.2400
Scaling   :   17.3913      0.2362      0.2300      0.2600
Summing   :   19.3549      0.3223      0.3100      0.3500
SAXPYing  :   20.0000      0.3095      0.3000      0.3600
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     93.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1819      0.2220      0.2200      0.2300
Scaling   :   17.3913      0.2341      0.2300      0.2400
Summing   :   19.3549      0.3140      0.3100      0.3200
SAXPYing  :   20.0000      0.3010      0.3000      0.3100
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     93.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   19.0476      0.2221      0.2100      0.2400
Scaling   :   17.3913      0.2330      0.2300      0.2400
Summing   :   19.3549      0.3151      0.3100      0.3300
SAXPYing  :   20.6897      0.3041      0.2900      0.3200
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     95.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1818      0.2200      0.2200      0.2200
Scaling   :   17.3913      0.2351      0.2300      0.2400
Summing   :   19.3548      0.3171      0.3100      0.3400
SAXPYing  :   20.0000      0.3041      0.3000      0.3200
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     94.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1818      0.2220      0.2200      0.2300
Scaling   :   17.3913      0.2330      0.2300      0.2400
Summing   :   19.3548      0.3140      0.3100      0.3200
SAXPYing  :   20.0000      0.3010      0.3000      0.3100

From Greg.Limes@Eng.Sun.COM  Tue Oct  1 22:01:56 1991
Received: from sun.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02161; Tue, 1 Oct 91 22:01:56 EDT
Received: from Eng.Sun.COM (zigzag-bb.Corp.Sun.COM) by Sun.COM (4.1/SMI-4.1)
	id AA04374; Tue, 1 Oct 91 19:00:26 PDT
Received: from ouroborous.Eng.Sun.COM by Eng.Sun.COM (4.1/SMI-4.1)
	id AA22070; Tue, 1 Oct 91 19:00:16 PDT
Received: by ouroborous.Eng.Sun.COM (4.1/SMI-4.1)
	id AA00755; Tue, 1 Oct 91 18:59:34 PDT
Date: Tue, 1 Oct 91 18:59:34 PDT
From: Greg.Limes@Eng.Sun.COM (Greg Limes)
Message-Id: <9110020159.AA00755@ouroborous.Eng.Sun.COM>
To: mccalpin
Subject: Memory bandwidth testing
Cc: Keith.Bierman@Eng.Sun.COM, Greg.Limes@Eng.Sun.COM
Content-Type: X-sun-attachment
Status: RO

----------
X-Sun-Data-Type: text
X-Sun-Data-Description: text
X-Sun-Data-Name: text
X-Sun-Content-Lines: 39


John -- I just happen to have a SPARCserver 600mp handy (ok, it's on my
desk, but that's handy :-) and ran your latest stuff. Here's my
results.

(Hope you don't mind the inclusion format, it is becoming fairly
standard and my tools do it for me automagicly)

I ran all this on a four-processor SPARCstation 670 (whatcha bet
marketing never quite calls it that) with 64Meg of memory, and all the
data of interest on an Elite 1.3gig SCSI disk.

To make the parallel runs more interesting, I bumped the repcount on
the loops by ten. Basicly, I want to be sure to get extended times
where all CPUs are banging on the problem, since I know that the system
is relatively benevolent to the bus and that is not really what you
want to measure.

MANIFEST
	out.Sun_670_s	single precision, one copy
	out.Sun_670_d	double precision, one copy
	out.Sun_670_2_s	single precision, two copies
	out.Sun_670_2_d	double precision, two copies
	out.Sun_670_4_s	single precision, four copies
	out.Sun_670_4_d	double precision, four copies

	Makefile	how I built the tests and made the Table
	runall		how I ran the bugger
	two.c		utility program used in runall
	four.c		utility program used in runall

	Table.awk	awk script used in making Table
	Table.ed	ed script used in making Table

	Table		nicely formatted summary of the data so far

Note: "Table.awk" assumes all runs in a file happened in parallel, and
adds up the rates to get a total rate. This may not be the right way to
calculate it. I'm a programmer, not a benchmarker ...
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: out.Sun_670_s
X-Sun-Content-Lines: 13

--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     110.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   21.0532      0.2022      0.1900      0.2500
Scaling   :   21.0528      0.2022      0.1900      0.2300
Summing   :   22.2223      0.2836      0.2700      0.3200
SAXPYing  :   24.0000      0.2586      0.2500      0.3100
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: out.Sun_670_d
X-Sun-Content-Lines: 13

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     145.99999040365 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2356      0.1789      0.1700      0.2300
Scaling   :   31.9997      0.1627      0.1500      0.2200
Summing   :   30.0003      0.2550      0.2400      0.3200
SAXPYing  :   29.9998      0.2532      0.2400      0.3000
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: out.Sun_670_2_s
X-Sun-Content-Lines: 26

--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     190.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   20.0000      0.2171      0.2000      0.2400
Scaling   :   19.0477      0.2236      0.2100      0.2800
Summing   :   20.6899      0.3042      0.2900      0.3400
SAXPYing  :   22.2225      0.2844      0.2700      0.3200
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     190.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   19.0477      0.2166      0.2100      0.2300
Scaling   :   19.0477      0.2245      0.2100      0.2800
Summing   :   20.6899      0.3037      0.2900      0.3900
SAXPYing  :   22.2225      0.2906      0.2700      0.6300
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: out.Sun_670_2_d
X-Sun-Content-Lines: 26

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     254.00001034141 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   22.8572      0.2164      0.2100      0.2500
Scaling   :   24.0004      0.2146      0.2000      0.2700
Summing   :   24.8282      0.3028      0.2900      0.3700
SAXPYing  :   25.7144      0.3008      0.2800      0.3700
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     247.99999073148 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   22.8572      0.2177      0.2100      0.2400
Scaling   :   24.0004      0.2128      0.2000      0.2700
Summing   :   24.8282      0.3023      0.2900      0.3700
SAXPYing  :   24.8282      0.3070      0.2900      0.6600
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: out.Sun_670_4_s
X-Sun-Content-Lines: 52

--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     246.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   14.2858      0.2979      0.2800      0.3600
Scaling   :   14.2858      0.2999      0.2800      0.3600
Summing   :   15.7899      0.4048      0.3800      0.4800
SAXPYing  :   16.6666      0.3845      0.3600      0.4500
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     285.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   14.8150      0.2959      0.2700      0.3600
Scaling   :   14.2858      0.2998      0.2800      0.5800
Summing   :   16.2162      0.4080      0.3700      0.4700
SAXPYing  :   17.1433      0.3857      0.3500      0.4300
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     314.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   14.2858      0.2982      0.2800      0.5400
Scaling   :   14.8154      0.2966      0.2700      0.3500
Summing   :   16.6666      0.4020      0.3600      0.4500
SAXPYing  :   16.6666      0.3914      0.3600      0.6300
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     317.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   19.9988      0.2944      0.2000      0.3800
Scaling   :   19.0484      0.2951      0.2100      0.3400
Summing   :   20.6890      0.4168      0.2900      0.9000
SAXPYing  :   22.2231      0.3869      0.2700      0.4500
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: out.Sun_670_4_d
X-Sun-Content-Lines: 52

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     425.00000447035 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   16.0007      0.3297      0.3000      0.8500
Scaling   :   16.5521      0.3089      0.2900      0.3600
Summing   :   16.7442      0.4671      0.4300      0.5300
SAXPYing  :   16.7442      0.4670      0.4300      0.5300
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     406.00000992417 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   16.0002      0.3232      0.3000      0.5000
Scaling   :   16.5521      0.3233      0.2900      0.7600
Summing   :   16.3641      0.4718      0.4400      0.8000
SAXPYing  :   18.4616      0.4609      0.3900      0.5400
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     464.99997973442 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   21.8181      0.3206      0.2200      0.3800
Scaling   :   22.8564      0.3105      0.2100      0.3600
Summing   :   22.5006      0.4664      0.3200      0.5200
SAXPYing  :   22.4995      0.4704      0.3200      0.7200
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     407.00000599027 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   25.2628      0.3296      0.1900      0.8600
Scaling   :   30.0022      0.3140      0.1600      0.5200
Summing   :   27.6913      0.4660      0.2600      0.6600
SAXPYing  :   27.6929      0.4723      0.2600      0.7700
----------
X-Sun-Data-Type: Makefile
X-Sun-Data-Description: Makefile
X-Sun-Data-Name: Makefile
X-Sun-Content-Lines: 20

.KEEP_STATE:

CC=	/usr/lang/cc
FC=	/usr/lang/f77
CFLAGS=	-fast -O4
FFLAGS=	-fast -O4
LDFLAGS=-Bstatic

PGMS=	stream_d stream_s two four

all: $(PGMS) Table

clean:
	-/bin/rm -f *.o $(PGMS)

Table: .FORCE
	grep . Results/* | sed 's;Results/;;' | sed 's/out.//' | sed 's/:/ /g' | awk -f Table.awk | sort +5 -nr > Table
	ed - Table < Table.ed

.FORCE:
----------
X-Sun-Data-Type: cshell-script
X-Sun-Data-Description: cshell-script
X-Sun-Data-Name: runall
X-Sun-Content-Lines: 7

#! /bin/csh -f
stream_d > Results/out.Sun_670_d
stream_s > Results/out.Sun_670_s
two stream_d > Results/out.Sun_670_2_d
two stream_s > Results/out.Sun_670_2_s
four stream_d > Results/out.Sun_670_4_d
four stream_s > Results/out.Sun_670_4_s
----------
X-Sun-Data-Type: c-file
X-Sun-Data-Description: c-file
X-Sun-Data-Name: two.c
X-Sun-Content-Lines: 10

main(ac,av)
	int ac;
	char **av;
{
    if (!*++av) exit(1);
    for (ac=0; ac<2; ++ac)
	if (!vfork())
	    exit(execvp(*av, av));
    while(wait((int *)0) != -1);
}
----------
X-Sun-Data-Type: c-file
X-Sun-Data-Description: c-file
X-Sun-Data-Name: four.c
X-Sun-Content-Lines: 10

main(ac,av)
	int ac;
	char **av;
{
    if (!*++av) exit(1);
    for (ac=0; ac<4; ++ac)
	if (!vfork())
	    exit(execvp(*av, av));
    while(wait((int *)0) != -1);
}
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: Table.awk
X-Sun-Content-Lines: 22

$3=="calibration"	{
	cal[$1] += $7;
	next;
	}
$2=="Assignment"	{
	asg[$1] += $3;
	}
$2=="Scaling"	{
	sca[$1] += $3;
	}
$2=="Summing"	{
	sum[$1] += $3;
	}
$2=="SAXPYing"	{
	sax[$1] += $3;
	}
END	{
		printf "%-20s %7s %7s %7s %7s %7s\n", "system", "cal", "copy", "scale", "sum", "saxpy";
		for (h in cal) {
			printf "%-20s %7.2f %7.1f %7.1f %7.1f %7.1f\n", h, cal[h], asg[h], sca[h], sum[h], sax[h];
		}
	}
----------
X-Sun-Data-Type: default
X-Sun-Data-Description: default
X-Sun-Data-Name: Table.ed
X-Sun-Content-Lines: 14

$m0
1a

.
$a

.
g/_s/m$
$a

.
g/_d/m$
w
q

From csrcb@shark.mel.dit.csiro.au  Tue Oct  1 22:02:29 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02165; Tue, 1 Oct 91 22:02:29 EDT
Received: by shark.mel.dit.csiro.au id AA03666
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Wed, 2 Oct 91 12:00:52 +1000
Date: Wed, 2 Oct 91 12:00:52 +1000
From: Robert Bell <Robert.Bell@mel.dit.csiro.au>
Message-Id: <9110020200.AA03666@shark.mel.dit.csiro.au>
To: mccalpin
Subject: re: memory bandwidth
Status: RO

John,
     Thanks for your comments.
     The C-90 has two pipes, and two load paths, one store path and one i/o
 path per processor.  This gives a peak of 8000 Mbyte/s for the assignment,
 and 12000 Mbyte/s for the operations with two vector reads and one vector
 store.
     I would love to work up the benchmark I have to use multiple processors.
 It needs a lot of coordination, because I would really like to have control
 over the exact cycle that each processor hits each bank.  I'll do some 
 preliminary tests.
     I think the C-90 and Y-MP can maintain that performance for each CPU
 provided you don't hit bank conflicts.  The machine I tested (in April)
 had only two processors, and the memory was by no means fully wired.  There
 was talk of something like 256 banks, compared with 64 for our Y-MP.
     Nostalgia - our organization has had a Cyber 205, Cyber 76, CDC 3600,
 CDC 3200, etc, and was looking at the ETAs at the time CDC made a mess of it.
 It was sad.  The 205 and ETA really knew how to provide fast memory transfers.
     Regards,
             Rob. Bell.

From csrcb@shark.mel.dit.csiro.au  Tue Oct  1 22:09:38 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02170; Tue, 1 Oct 91 22:09:38 EDT
Received: by shark.mel.dit.csiro.au id AA03839
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Wed, 2 Oct 91 12:08:06 +1000
Date: Wed, 2 Oct 91 12:08:06 +1000
From: Robert Bell <Robert.Bell@mel.dit.csiro.au>
Message-Id: <9110020208.AA03839@shark.mel.dit.csiro.au>
To: mccalpin
Subject: Re:  Benchmark
Status: RO

 John,
      Thanks for the hints about multitasking - I'll try my benchmark again.
      Rob.

From @scapa.cs.ualberta.ca:kufeld_k@edhp01  Tue Oct  1 22:58:25 1991
Received: from scapa.cs.ualberta.ca by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02181; Tue, 1 Oct 91 22:58:25 EDT
Received: from ipl by scapa.cs.ualberta.ca with UUCP id <42586>; Tue, 1 Oct 1991 20:56:45 -0600
Received: by edhp01
	(16.7/16.2) id AA02938; Tue, 1 Oct 91 20:56:44 -0600
Date: 	Tue, 1 Oct 1991 20:56:44 -0600
From: Kurt Kufeld <kufeld_k@edhp01.uucp>
To: mccalpin
Subject: HP9000/730 Memory Timings
Message-Id: <91Oct1.205645mdt.42586@scapa.cs.ualberta.ca>
Status: RO

John:

	Here are the results of the memory bandwidth test on a 730.  I was
unable to run them even with "n" set to 1 000 000 and 500 000 set for single
and double precision.  I had to reduce the value of "n" to 600 000 and 300 000
for single and double respectively.  It coredumped otherwise.

	Lucky one of our systems has a Fortran compiler.

	If these test are good enough, let me know.  If there is some way
to get them to run with a larger value of n let me know.  I did not dig deeply
into the code.

	Here are the specifics:

		Hardware:	HP9000/730,  32 Mbytes of memory
		O.S.:		HP-UX 08.05
		Compilers:
			C:	HP C/ANSI C compiler
			flags:	+O3 -osecond
			Fortran:HP fortran
			flags:	-O -sstream_x

	If I'm missing anything, let me know.

	Let me know how we compare.  I missed your initial comparisons.


--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =  27.00000014156103 hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   53.3334       .0951       .0900       .1000  
Scaling   :   48.0000       .1000       .1000       .1000  
Summing   :   55.3847       .1300       .1300       .1300  
SAXPYing  :   55.3848       .1381       .1300       .1400  


--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time =  31.0 hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   40.0000       .1231       .1200       .1300  
Scaling   :   36.9231       .1310       .1300       .1400  
Summing   :   45.0000       .1661       .1600       .1700  
SAXPYing  :   40.0000       .1841       .1800       .1900  

*****************************************************************************

					Kurt Kufeld
					kufeld_k%ipl@cs.ualberta.ca
					(403) 420-5330

From csrcb@shark.mel.dit.csiro.au  Wed Oct  2 10:00:08 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02915; Wed, 2 Oct 91 10:00:08 EDT
Received: by shark.mel.dit.csiro.au id AA09657
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Wed, 2 Oct 91 23:58:39 +1000
Date: Wed, 2 Oct 91 23:58:39 +1000
From: Robert Bell <Robert.Bell@mel.dit.csiro.au>
Message-Id: <9110021358.AA09657@shark.mel.dit.csiro.au>
To: mccalpin
Subject: re: memory bandwidth
Status: RO

 John,
      Each memory port can move two words per cycle, corresponding to the two 
pipes.  I don't know in the hardware whether it is that there are two pipes
in each of the four ports, or four ports in each pipe, but the effect is the 
same.  So for the triad, the C-90 theoretical peak is
 (2 read +1 write) * 2 pipes * 8 bytes * 16 processors / 4.0e-9 = 192 Gbytes/s.
 Similarly for the Y-MP8 (2 +1) * 8 bytes * 8 / 6.0e-9 = 32 Gbytes/s.
 Andrew Zachary's results of 26.802 Gbyte/s for the Y-MP8 are consistent
 with that - the loss is because of vector startup and bank conflicts, etc.

      I did the autotasking tests on our Y-MP2, but could not get consistent
 results because the machine was busy (measuring elapsed rather than CPU time).
      I think the next test to do is to check the results with the arrays
 aligned differently, i.e. a(i) = b(i+k) for various k - my results show a
 factor of two difference with different k's.  Might I suggest making the arrays
 of length 2**20, and putting them in common so that there is a consistent
 alignment of the start of each array in the same bank.

      Cheers,
             Rob. Bell.

From UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper  Wed Oct  2 12:58:12 1991
Received: from ucsd.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03134; Wed, 2 Oct 91 12:58:12 EDT
Received: from celit.UUCP by ucsd.edu; id AA21092
	sendmail 5.64/UCSD-2.2-sun via UUCP
	Wed, 2 Oct 91 09:44:12 -0700
Received: by celit.fps.com (5.51/celerity1.1)
	id AA11756; Wed, 2 Oct 91 09:40:10 PDT for mccalpin@perelandra.cms.udel.edu at ucsd
Posted-Date: Wed, 2 Oct 91 09:39:10 PDT
Received: from keeper.fps_net by fpssun (4.1/SMI-4.1)
	id AA18886; Wed, 2 Oct 91 09:32:44 PDT
Received: by keeper.fps_net (4.1/SMI-4.1)
	id AA06864; Wed, 2 Oct 91 09:39:10 PDT
Date: Wed, 2 Oct 91 09:39:10 PDT
From: UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper (Brian Whitney)
Message-Id: <9110021639.AA06864@keeper.fps_net>
To: mccalpin
Subject: More memory bandwidth testing
Status: RO


Hi, 

I have found another interesting thing compilers can do to your
new test code and a simple workaround for it.

For some reason (I need to figure out what the compiler did) our
(FPS) compiler interchanged the 60 loop and invalidated the 
timings generated.  I saw megabytes beyond what the machine is
capable of achieving.  So I modified the code  to look like this

      DO 60 k = 1,ntimes

          call dummy2(a,b,c)
          t = second(t0)
          DO 20 j = 1,n
              c(j) = a(j)
   20     CONTINUE
          t = second(t0) - t
          times(1,k) = t

          call dummy2(a,b,c)
          t = second(t0)
          DO 30 j = 1,n
              c(j) = 3.0D0*a(j)
   30     CONTINUE
          t = second(t0) - t
          times(2,k) = t

          call dummy2(a,b,c)
          t = second(t0)
          DO 40 j = 1,n
              c(j) = a(j) + b(j)
   40     CONTINUE
          t = second(t0) - t
          times(3,k) = t

          call dummy2(a,b,c)
          t = second(t0)
          DO 50 j = 1,n
              c(j) = a(j) + 3.0D0*b(j)
   50     CONTINUE
          t = second(t0) - t
          times(4,k) = t
   60 CONTINUE

where dummy2 is

      SUBROUTINE dummy2(q,r,s)
C     .. Scalar Arguments ..
      DOUBLE PRECISION q,r,s
C     .. 
      RETURN
      END

This caused the compiler to assume I messed with a, b, c before each
call and allowed proper execution.

I am still gather data.  I am running on about 10 different types of 
machines, so I want to make sure the data is reasonable.  (And I am
doing this "on my own", while waiting for compiles.)

Brian Whitney
FPS Computing

keeper@fps.com

From Keith.Bierman@Eng.Sun.COM  Wed Oct  2 13:08:12 1991
Received: from sun.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03157; Wed, 2 Oct 91 13:08:12 EDT
Received: from Eng.Sun.COM (zigzag-bb.Corp.Sun.COM) by Sun.COM (4.1/SMI-4.1)
	id AA21516; Wed, 2 Oct 91 10:06:46 PDT
Received: from chiba.Eng.Sun.COM by Eng.Sun.COM (4.1/SMI-4.1)
	id AA03196; Wed, 2 Oct 91 10:06:32 PDT
Received: by chiba.Eng.Sun.COM (4.1/SMI-4.1)
	id AA00495; Wed, 2 Oct 91 10:06:27 PDT
Date: Wed, 2 Oct 91 10:06:27 PDT
From: Keith.Bierman@Eng.Sun.COM (Keith Bierman fpgroup)
Message-Id: <9110021706.AA00495@chiba.Eng.Sun.COM>
To: mccalpin (John D. McCalpin)
In-Reply-To: mccalpin@perelandra.cms.udel.edu's message of 2 Oct 91 13:43:37 GMT
Subject: Attainable Memory Bandwidth (update)
Status: RO


Looks like my SS-2 fell off the chart ;<

From preston@rice.edu  Wed Oct  2 13:16:31 1991
Received: from rice.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03162; Wed, 2 Oct 91 13:16:31 EDT
Received: from dawn.rice.edu by rice.edu (AA02045); Wed, 2 Oct 91 12:14:30 CDT
Received: by dawn.rice.edu (AA08648); Wed, 2 Oct 91 12:14:57 CDT
Date: Wed, 2 Oct 91 12:14:57 CDT
From: preston@rice.edu (Preston Briggs)
Message-Id: <9110021714.AA08648@dawn.rice.edu>
To: mccalpin
Subject: Re: Attainable Memory Bandwidth (update)
Newsgroups: comp.arch
In-Reply-To: <MCCALPIN.91Oct2094337@pereland.cms.udel.edu>
Organization: Rice University, Houston
Cc: 
Status: RO


This was super interesting to me.
Thanks for taking the time.

Regards,
Preston

From dave@hanauma.stanford.edu  Wed Oct  2 13:43:44 1991
Received: from oas.Stanford.EDU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03182; Wed, 2 Oct 91 13:43:44 EDT
Received: from hanauma.stanford.edu by oas.Stanford.EDU (4.1/inc-1.0)
	id AA19224; Wed, 2 Oct 91 10:42:19 PDT
Received:  by hanauma.stanford.edu (5.64/7.0) 
		 id AA18594; Wed, 2 Oct 91 10:42:17 -0700
Date:  Wed, 2 Oct 91 10:42:17 -0700
From: dave@hanauma.STANFORD.EDU (Dave Nichols)
Message-Id:  <9110021742.AA18594@hanauma.stanford.edu>
To: mccalpin
Subject: old convex mem speed
Status: RO


These times are from a moderately loaded convex C-1XP (1985 vintage)

--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    88.60040     hundredths of a second
Increase the size of the arrays if this is <30
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   62.9030      0.1331      0.1272      0.1382
Scaling   :   62.5127      0.1312      0.1280      0.1382
Summing   :   64.4268      0.1938      0.1863      0.2067
SAXPYing  :   64.4773      0.1926      0.1861      0.2063

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =    155.416792631149      hundredths of a second
Increase the size of the arrays if this is <30
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   71.2970      0.2100      0.2020      0.2328
Scaling   :   71.0304      0.2121      0.2027      0.2229
Summing   :   71.2635      0.3113      0.3031      0.3247
SAXPYing  :   71.3151      0.3130      0.3029      0.3230


Dave Nichols, Dept. of Geophysics, Stanford University.
dave@hanauma.stanford.edu

From klee@nas.nasa.gov  Wed Oct  2 14:23:08 1991
Received: from wilbur.nas.nasa.gov by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03194; Wed, 2 Oct 91 14:23:08 EDT
Received: by wilbur.nas.nasa.gov (5.61/1.34)
	id AA07974; Wed, 2 Oct 91 11:21:42 -0700
Date: Wed, 2 Oct 91 11:21:42 -0700
From: klee@nas.nasa.gov (King M. Lee)
Message-Id: <9110021821.AA07974@wilbur.nas.nasa.gov>
To: mccalpin
Subject: Re: Attainable Memory Bandwidth (update)
Newsgroups: comp.arch
In-Reply-To: <MCCALPIN.91Oct2094337@pereland.cms.udel.edu>
Organization: NASA Ames Research Center, Moffett Field, CA
Cc: klee@nas.nasa.gov
Status: RO

I enjoyed your posting on attainable memory bandwidth.  I noticed
you had nothing on i860.  In the summer of 1990, I did some work
on the i860.  I had a DCOPY routine that effectively  read and wrote
data at the rate of 12.6 million double precision words per second
which works out to about 100 Mbytes second.  This was done in
assembler using pipelined loads.  The results are in
"On the Floating Point performance of the i860(TM) 
Miicroprocessor", by King Lee, report RNR-90-019; NAS Systems 
Division; Ames Research Center; Moffett field. CA 94035.
I  can send you a copy if you want. 

I am away from Ames now, and it is not convenient to run you 
program.   If I get a chance I'll run your program and send you
the results.

King Lee



From Greg.Limes@Eng.Sun.COM  Wed Oct  2 15:22:21 1991
Received: from sun.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03216; Wed, 2 Oct 91 15:22:21 EDT
Received: from Eng.Sun.COM (zigzag-bb.Corp.Sun.COM) by Sun.COM (4.1/SMI-4.1)
	id AA16399; Wed, 2 Oct 91 12:20:55 PDT
Received: from ouroborous.Eng.Sun.COM by Eng.Sun.COM (4.1/SMI-4.1)
	id AA11115; Wed, 2 Oct 91 12:20:43 PDT
Received: by ouroborous.Eng.Sun.COM (4.1/SMI-4.1)
	id AA01716; Wed, 2 Oct 91 12:20:02 PDT
Date: Wed, 2 Oct 91 12:20:02 PDT
From: Greg.Limes@Eng.Sun.COM (Greg Limes)
Message-Id: <9110021920.AA01716@ouroborous.Eng.Sun.COM>
To: mccalpin
Subject: Re:  Memory bandwidth testing
Status: RO


John -- glad to be of help. I think I am not leaking any secrets if I
say that the CPUs are SPARC chips running at 40 mhz. If you want more
details about the system, tho, I'll have to ask what bits of information
we are stamping as trade secret; most of the interesting things should
be easily public.

From Keith.Bierman@Eng.Sun.COM  Wed Oct  2 16:47:37 1991
Received: from sun.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03371; Wed, 2 Oct 91 16:47:37 EDT
Received: from Eng.Sun.COM (zigzag-bb.Corp.Sun.COM) by Sun.COM (4.1/SMI-4.1)
	id AA00956; Wed, 2 Oct 91 13:46:13 PDT
Received: from chiba.Eng.Sun.COM by Eng.Sun.COM (4.1/SMI-4.1)
	id AA14249; Wed, 2 Oct 91 13:46:00 PDT
Received: from localhost by chiba.Eng.Sun.COM (4.1/SMI-4.1)
	id AA01488; Wed, 2 Oct 91 13:45:55 PDT
Message-Id: <9110022045.AA01488@chiba.Eng.Sun.COM>
To: "John D. McCalpin" <mccalpin>
Subject: Re: Attainable Memory Bandwidth (update) 
In-Reply-To: Your message of Wed, 02 Oct 91 16:22:55 -0400.
             <9110022022.AA03304@perelandra.cms.udel.edu> 
Date: Wed, 02 Oct 91 13:45:54 PDT
From: Keith.Bierman@Eng.Sun.COM
Status: RO


chiba, a 4/75 (aka SPARCstation 2), 40mhz, etc. sunos 4.1.1, compiler
options (v1.4) -fast -O4 -Bstatic, OW3.0 running.

real*16 was too slow to be of interest ;<

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     99.999996274710 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2354      0.1791      0.1700      0.1900
Scaling   :   30.0000      0.1696      0.1600      0.2100
Summing   :   28.8000      0.2571      0.2500      0.2700
SAXPYing  :   26.6667      0.2791      0.2700      0.3000
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     97.999998182058 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   30.0000      0.1711      0.1600      0.1800
Scaling   :   30.0000      0.1661      0.1600      0.1800
Summing   :   28.8000      0.2530      0.2500      0.2600
SAXPYing  :   26.6667      0.2740      0.2700      0.2800
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     97.999998182058 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2353      0.1814      0.1700      0.2000
Scaling   :   30.0000      0.1762      0.1600      0.1900
Summing   :   28.8000      0.2694      0.2500      0.2900
SAXPYing  :   26.6667      0.2925      0.2700      0.3200
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =    104.999991506338 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2354      0.1843      0.1700      0.2000
Scaling   :   30.0000      0.1773      0.1600      0.2000
Summing   :   28.8000      0.2695      0.2500      0.2900
SAXPYing  :   26.6667      0.2885      0.2700      0.3200
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     107.00000151992 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   28.2353      0.1824      0.1700      0.2000
Scaling   :   30.0000      0.1793      0.1600      0.2000
Summing   :   28.8000      0.2672      0.2500      0.2800
SAXPYing  :   26.6666      0.2883      0.2700      0.3100
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     95.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1819      0.2251      0.2200      0.2400
Scaling   :   17.3913      0.2362      0.2300      0.2600
Summing   :   19.3549      0.3223      0.3100      0.3500
SAXPYing  :   20.0000      0.3095      0.3000      0.3600
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     93.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1819      0.2220      0.2200      0.2300
Scaling   :   17.3913      0.2341      0.2300      0.2400
Summing   :   19.3549      0.3140      0.3100      0.3200
SAXPYing  :   20.0000      0.3010      0.3000      0.3100
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     93.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   19.0476      0.2221      0.2100      0.2400
Scaling   :   17.3913      0.2330      0.2300      0.2400
Summing   :   19.3549      0.3151      0.3100      0.3300
SAXPYing  :   20.6897      0.3041      0.2900      0.3200
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     95.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1818      0.2200      0.2200      0.2200
Scaling   :   17.3913      0.2351      0.2300      0.2400
Summing   :   19.3548      0.3171      0.3100      0.3400
SAXPYing  :   20.0000      0.3041      0.3000      0.3200
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     94.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   18.1818      0.2220      0.2200      0.2300
Scaling   :   17.3913      0.2330      0.2300      0.2400
Summing   :   19.3548      0.3140      0.3100      0.3200
SAXPYing  :   20.0000      0.3010      0.3000      0.3100

From uunet.UU.NET!stardent!Stardent.COM!wright  Wed Oct  2 16:53:30 1991
Received: from relay1.UU.NET by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03399; Wed, 2 Oct 91 16:53:30 EDT
Received: from uunet.uu.net (via LOCALHOST.UU.NET) by relay1.UU.NET with SMTP 
	(5.61/UUNET-internet-primary) id AA13274; Wed, 2 Oct 91 16:52:06 -0400
Received: from stardent.UUCP by uunet.uu.net with UUCP/RMAIL
	(queueing-rmail) id 165141.26558; Wed, 2 Oct 1991 16:51:41 EDT
Received: by stardent.Stardent.COM (1.1/smail2.5/01-28-89)
	id AA00449; Wed, 2 Oct 91 16:46:41 EDT
Message-Id: <9110022046.AA00449@stardent.Stardent.COM>
To: uunet.UU.NET!uunet!perelandra.cms.udel.edu!mccalpin
Cc: mdavis@Stardent.COM, swin@Stardent.COM
Subject: Re: Attainable Memory Bandwidth
Date: Wed, 02 Oct 91 16:46:39 EDT
From: David Wright <wright@Stardent.COM>
Status: RO

[I hope you haven't made your last posting on this subject, since
there might well be other contributors; I got back as quickly as I
could.] 

Here are the results of your benchmark for two Stardent machines,
the Vistra 800b and the ST2000.

(This is the double-precision benchmark in all cases.)

1)  Vistra 800b (i860-based)  Compiler:  Portland Group FTN, Rev 1.4 (beta)

  Compiler output:

    f77 -O4 -Mvect -Mbeta -Mx,0,2 -o mcc2 mcc2.f

    PGFTN-I-Beta Release Optimizations Activated
    Vect: streaming data and stripmining loop at line 106. strip size = 252.
    Vect: loop at line 106 replaced by call to __add8s.
    Vect: streaming data and stripmining loop at line 99. strip size = 252.
    Vect: loop at line 99 replaced by call to __add8s.
    Vect: streaming data and stripmining loop at line 92. strip size = 504.
    Vect: streaming data and stripmining loop at line 85. strip size = 504.
    Vect: streaming data and stripmining loop at line 69. strip size = 252.
    # SW pipelined loop w/ 23 cycles and 2 columns w/ cnt 4 gend for line 115
    Linking:

  Runtime output:

    --------------------------------------
     Double precision appears to have 16 digits of accuracy
     Assuming 8 bytes per DOUBLEPRECISION word
    --------------------------------------
     Timing calibration ; time = 99.99999627470970 hundredths of a second
     Increase the size of the arrays if this is <30
      and your clock precision is =<1/100 second
     ---------------------------------------------------
    Function     Rate (MB/s)  RMS time   Min time  Max time
    Assignment:  160.0002      0.0373      0.0300      0.0400
    Scaling   :  160.0014      0.0354      0.0300      0.0400
    Summing   :  120.0001      0.0652      0.0600      0.0700
    SAXPYing  :  120.0001      0.0652      0.0600      0.0700


2)  Stardent ST2000  Compiler:  Stardent f77 compiler (Rev 2.3)
    (second() modified for Stellix) (vector parallel)

  Compiler output:

    f77 -O3 -o mcc2 mcc2_gs.f

    "mcc2_gs.f", line 125: Advisory: argument variable 'RMSTIME' may inhibit 
	vectorization.
    "mcc2_gs.f", line 125: Advisory: argument variable 'MINTIME' may inhibit
	vectorization.
    "mcc2_gs.f", line 125: Advisory: argument variable 'MAXTIME' may inhibit
	vectorization.
    "mcc2_gs.f", line 69: Loop (J-loop) fully parallelized
    "mcc2_gs.f", line 69: Loop (J-loop) fully vectorized
    "mcc2_gs.f", line 82: Loop (K-loop) contains a subroutine or function 
	call or character expressions
    "mcc2_gs.f", line 82: Loop (K-loop) not vectorized
    "mcc2_gs.f", line 85: Loop (J-loop) fully parallelized
    "mcc2_gs.f", line 85: Loop (J-loop) fully vectorized
    "mcc2_gs.f", line 92: Loop (J-loop) fully parallelized
    "mcc2_gs.f", line 92: Loop (J-loop) fully vectorized
    "mcc2_gs.f", line 99: Loop (J-loop) fully parallelized
    "mcc2_gs.f", line 99: Loop (J-loop) fully vectorized
    "mcc2_gs.f", line 106: Loop (J-loop) fully parallelized
    "mcc2_gs.f", line 106: Loop (J-loop) fully vectorized
    "mcc2_gs.f", line 115: Loop (J-loop) fully vectorized
    "mcc2_gs.f", line 122: Loop (J-loop) performs I/O or contains character 
	expressions
    "mcc2_gs.f", line 122: Loop (J-loop) not vectorized

            Vectorization-Parallelization Summary for Routine STREAM

     Line Index  Label  Start   Stop Step  Vec.  Par.  Reason
    ---------------------------------------------------------------------------
       69 J         10      1 300000    1  FULL  FULL
       82 K         60      1     10    1  None  None  Library or function call
       85 J         20      1 300000    1  FULL  FULL
       92 J         30      1 300000    1  FULL  FULL
       99 J         40      1 300000    1  FULL  FULL
      106 J         50      1 300000    1  FULL  FULL
      114 K         80      1     10    1  None  None  Outer loop of nest
      115 J         70      1      4    1  FULL  None
      122 J         90      1      4    1  None  None  I/O or char expressions

    "mcc2_gs.f", line 181: Warning: label '10' defined but not referenced.
    "mcc2_gs.f", line 181: Loop (J-loop) fully vectorized
    "mcc2_gs.f", line 185: Loop (J-loop) or a contained loop has multiple exits
    "mcc2_gs.f", line 185: Loop (J-loop) not vectorized

            Vectorization-Parallelization Summary for Routine REALSIZE

     Line Index  Label   Start   Stop  Step  Vec.  Par.  Reason
    ---------------------------------------------------------------------------
      181 J         20       1     30     1  FULL  None
      185 J         30       1     30     1  None  None  Inner loop: many exits

  Runtime output:

  --------------------------------------
   Double precision appears to have 16 digits of accuracy
   Assuming 8 bytes per DOUBLEPRECISION word
  --------------------------------------
  Timing calibration ; time =    72.9492187500000      hundredths of a second
  Increase the size of the arrays if this is <30
   and your clock precision is =<1/100 second
  ---------------------------------------------------
  Function     Rate (MB/s)  RMS time   Min time  Max time
  Assignment:  163.8400      0.0357      0.0293      0.0498
  Scaling   :  163.8400      0.0353      0.0293      0.0400
  Summing   :  179.8244      0.0463      0.0400      0.0508
  SAXPYing  :  179.8244      0.0502      0.0400      0.0605

From kate@ahab.rutgers.edu  Wed Oct  2 17:22:29 1991
Received: from ahab.rutgers.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03447; Wed, 2 Oct 91 17:22:29 EDT
Message-Id: <9110022122.AA03447@perelandra.cms.udel.edu>
Received: by ahab.rutgers.edu id AA29243g; Wed, 2 Oct 91 17:22:05 EDT
Date: Wed, 2 Oct 91 17:22:05 EDT
From: Kate Hedstrom <kate@ahab.rutgers.edu>
To: mccalpin
Subject: Re:  memory bandwidth
Status: RO

John,

Sure, I'd be happy to.  I also have access to a Sparc 2 - should I run
it there too?

Kate

From klee@nas.nasa.gov  Wed Oct  2 18:10:50 1991
Received: from wilbur.nas.nasa.gov by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03515; Wed, 2 Oct 91 18:10:50 EDT
Received: by wilbur.nas.nasa.gov (5.61/1.34)
	id AA14342; Wed, 2 Oct 91 15:09:23 -0700
Date: Wed, 2 Oct 91 15:09:23 -0700
From: klee@nas.nasa.gov (King M. Lee)
Message-Id: <9110022209.AA14342@wilbur.nas.nasa.gov>
To: "John D. McCalpin" <mccalpin>
Subject: Re: Attainable Memory Bandwidth (update)
Status: RO

I also did a DCOPY using cached rather than pipelined loads (assembly).
I got about copied about 4.5 M double precision words, or
2 * 4.5 * 8 = 72 Mbytes per second.  I suspect that the compiled code
would get less, but I'm not  how much less it will get.  
I think compiled code gets worse for more complex loops.

King

From uunet.UU.NET!stardent!Stardent.COM!wright  Wed Oct  2 18:19:48 1991
Received: from relay1.UU.NET by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03521; Wed, 2 Oct 91 18:19:48 EDT
Received: from uunet.uu.net (via LOCALHOST.UU.NET) by relay1.UU.NET with SMTP 
	(5.61/UUNET-internet-primary) id AA08788; Wed, 2 Oct 91 18:18:25 -0400
Received: from stardent.UUCP by uunet.uu.net with UUCP/RMAIL
	(queueing-rmail) id 181744.24604; Wed, 2 Oct 1991 18:17:44 EDT
Received: by stardent.Stardent.COM (1.1/smail2.5/01-28-89)
	id AA03003; Wed, 2 Oct 91 18:01:59 EDT
Message-Id: <9110022201.AA03003@stardent.Stardent.COM>
To: uunet.UU.NET!uunet!perelandra.cms.udel.edu!mccalpin
Subject: Re: Attainable Memory Bandwidth
Date: Wed, 02 Oct 91 18:01:55 EDT
From: David Wright <wright@Stardent.COM>
Status: RO


>If these machines use the standard UNIX approach of measuring cpu time
>in 1/100's of a second, then these cases need to be run with much
>longer vectors.  They look like only 3-4 clock ticks for the fastest
>cases, and this leaves a 25-33% error in the estimates.

Spoilsport.  OK, I reran the tests with vectors 4x longer than the
original program.  These results are very repeatable and didn't vary
by more than 1% from run-to-run.  (Yes, the first run is apt to be
slower due to faulting in the pages.)

1) Vistra 800b

 Timing calibration ; time =    375.9999923408031      hundredths of a second
 Increase the size of the arrays if this is <30
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  147.6933      0.1331      0.1300      0.1400
Scaling   :  147.6933      0.1341      0.1300      0.1400
Summing   :  115.2000      0.2540      0.2500      0.2600
SAXPYing  :  115.2000      0.2560      0.2500      0.2600


2) ST2000:  no change needed.  The 2000 has a relatively slow clock
rate and I'd already tried it with vectors twice as long without
seeing any difference; doubling it again still made no difference
(although the results started to vary more due to more contention for
the cpu from other sources).

>By the way, what is the clock speed on the Vistra 800b?

40 MHz.  (FYI, this is also true for the V800e and V800ex, which are
different graphics, not different host architecture.)

>And is the ST2000 one of the older MIPS-based machines?

Nope.  That's the P2/P3 (now known as the ST1500 and ST3000).  The
ST2000 is the original Stellar architecture (proprietary).  Still
being sold, (including refurbished used machines, mostly the earlier
but very similar ST1000), and the price/perf isn't bad at all.

>I have access to a P3 that I will get this stuff run on, too....

Good.  I tried it myself here and got numbers in the ballpark of the
ST2000 (slightly better, maybe 15%), but since I'm an old Stellarite,
I wasn't sure I was using the compiler/system to best advantage.  (I
don't even know all the details on the ST2000 -- "Dammit, Jim, I'm a
kernel programmer, not a benchmarker!")

  -- David Wright, Stardent Computer Inc
     wright@stardent.com

From kate@ahab.rutgers.edu  Wed Oct  2 18:23:24 1991
Received: from ahab.rutgers.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03525; Wed, 2 Oct 91 18:23:24 EDT
Message-Id: <9110022223.AA03525@perelandra.cms.udel.edu>
Received: by ahab.rutgers.edu id AA29271g; Wed, 2 Oct 91 18:23:01 EDT
Date: Wed, 2 Oct 91 18:23:01 EDT
From: Kate Hedstrom <kate@ahab.rutgers.edu>
To: mccalpin
Subject: Re:  memory bandwidth
Status: RO

John,

We have etime, not second().  Here is what the man page says - is dtime
what I want for a multi-cpu test?

> SYNOPSIS
>      EXTERNAL ETIME
>      REAL ETIME
>      REAL TARRAY(2)
>      REAL TSECS
>      TSECS = ETIME(TARRAY)
> 
>      EXTERNAL DTIME
>      REAL DTIME
>      REAL TARRAY(2)
>      REAL DTSECS
>      DTSECS = DTIME(TARRAY)
> 
> DESCRIPTION
>      These two routines return elapsed runtime in seconds for the
>      calling process.  DTIME returns the elapsed time since the
>      last call to DTIME or the start of execution on the first
>      call.
> 
>      The argument array returns user time in the first element
>      and system time in the second element.  The function value
>      is the sum of user and system time.
> 
>      The resolution of all timing is 1/100th of a second.

The output:

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    749.0000307559967     hundredths
 of a second
 Increase the size of the arrays if this is <30  and your clock
 precision is =<1/100 second
---------------------------------------------------
 Function     Rate (MB/s)  RMS time   Min time  Max time
 Assignment:   83.1168      0.8135      0.7700      0.8700
 Scaling   :   78.0488      0.8808      0.8200      0.9700
 Summing   :   78.0488      1.3065      1.2300      1.3900
 SAXPYing  :   70.0730      1.4586      1.3700      1.5100

Kate

From kate@ahab.rutgers.edu  Wed Oct  2 18:28:36 1991
Received: from ahab.rutgers.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03530; Wed, 2 Oct 91 18:28:36 EDT
Message-Id: <9110022228.AA03530@perelandra.cms.udel.edu>
Received: by ahab.rutgers.edu id AA29302g; Wed, 2 Oct 91 18:28:12 EDT
Date: Wed, 2 Oct 91 18:28:12 EDT
From: Kate Hedstrom <kate@ahab.rutgers.edu>
To: mccalpin
Subject: Re:  memory bandwidth
Status: RO

John,

I just ran it again with 3 processors ( and 3 other jobs going):

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    768.9999878406525     hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  228.5716      0.3691      0.2800      0.4400
Scaling   :  246.1536      0.3468      0.2600      0.4100
Summing   :  246.1536      0.4819      0.3900      0.5500
SAXPYing  :  200.0002      0.5412      0.4800      0.6100

Kate

From csrcb@shark.mel.dit.csiro.au  Wed Oct  2 18:58:26 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03701; Wed, 2 Oct 91 18:58:26 EDT
Received: by shark.mel.dit.csiro.au id AA12981
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Thu, 3 Oct 91 08:57:00 +1000
Date: Thu, 3 Oct 91 08:57:00 +1000
From: Robert Bell <Robert.Bell@mel.dit.csiro.au>
Message-Id: <9110022257.AA12981@shark.mel.dit.csiro.au>
To: mccalpin
Subject: re: memory bandwidth
Status: RO

 John,
      I did mean but did not say that the results can be a factor of two worse
 with different alignments.  This happens on an X-MP, but not on the Y-MP, where
 the memory is better.
      It would be nice to add a column of bytes/cycle to your table of results.
 The C-90 theoretical peak is 3*2*8*16 = 768 bytes/cycle, or 1024 bytes/cycle
 if you include the i/o ports as well.
      The results from the Convexes will be interesting.  Since there is only 
 one port to memory for each CPU, the rate is 8 bytes/cycle per CPU, or
 32 bytes/cycle for C2 and 64 bytes/cycle for C3.  Thus the SAXPY operation
 can proceed only at one third of the speed the CPU is capable of, since the
 data transfer is the bottleneck.  The Linpack 1000 results are obtained by
 rewriting the code with much unrolling to have only one vector reference
 for each two arithmetic operations.
      I now have good autotasking results from our Y-MP2/216.
      Alignment and stride tests are fun!  You can make the Y-MP slow down by
 a factor of 15 (12 on later models) by striding all the vectors through with
 a stride of k*(number of banks).  On the Convexes (and I suspect the IBM 3090 
 VF), with large strides you hit cache problems in a big way.
      Cheers,
             Rob.
 P.S. Another useful figure would be peak memory transfer rate / peak meaflops,
 as a measure of the memory performance relative to what most people look at.

From cks@hawkwind.utcs.toronto.edu  Thu Oct  3 00:37:30 1991
Received: from hawkwind.utcs.utoronto.ca by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03931; Thu, 3 Oct 91 00:37:30 EDT
Received: from localhost by hawkwind.utcs.toronto.edu with SMTP id <2714>; Thu, 3 Oct 1991 00:35:57 -0400
To: "John D. McCalpin" <mccalpin>
Subject: Re: Attainable Memory Bandwidth (update)
Newsgroups: comp.arch
In-Reply-To: <MCCALPIN.91Oct2094337@pereland.cms.udel.edu>
Organization: Ziebmef home away from home
Date: 	Thu, 3 Oct 1991 00:35:40 -0400
From: Chris Siebenmann <cks@hawkwind.utcs.toronto.edu>
Message-Id: <91Oct3.003557edt.2714@hawkwind.utcs.toronto.edu>
Status: RO

 Some results, from a DECSystem 5000/200 with the DEC Fortran 3.0
Fortran compiler:

! From:	The Batch Daemon <batch@rw.cquest.toronto.edu>
! Subject: on server.rw: Success+Output: 'cfq246008106' in 'now' queue
! 
! Job 'cfq246008106' in 'now' queue has completed successfully.
! Total CPU used: 42.3 sec.
! Command list follows:
! 
! ./a.out; ./a.out; ./a.out; ./a.out
! 
! Output follows:
! 
! --------------------------------------
!  Double precision appears to have 16 digits of accuracy
!  Assuming 8 bytes per DOUBLEPRECISION word
! --------------------------------------
!  Timing calibration ; time =    50.38740262389183     hundredths of a second
!  Increase the size of the arrays if this is <30 
!   and your clock precision is =<1/100 second
!  ---------------------------------------------------
! Function     Rate (MB/s)  RMS time   Min time  Max time
! Assignment:   28.5786      0.1792      0.1680      0.1992
! Scaling   :   24.5776      0.2024      0.1953      0.2148
! Summing   :   24.2542      0.3131      0.2969      0.3320
! SAXPYing  :   23.0415      0.3302      0.3125      0.3672
! --------------------------------------
!  Double precision appears to have 16 digits of accuracy
!  Assuming 8 bytes per DOUBLEPRECISION word
! --------------------------------------
!  Timing calibration ; time =    49.99680016189814     hundredths of a second
!  Increase the size of the arrays if this is <30 
!   and your clock precision is =<1/100 second
!  ---------------------------------------------------
! Function     Rate (MB/s)  RMS time   Min time  Max time
! Assignment:   28.5786      0.1715      0.1680      0.1758
! Scaling   :   24.5776      0.1977      0.1953      0.2031
! Summing   :   24.2542      0.3008      0.2969      0.3047
! SAXPYing  :   23.0415      0.3172      0.3125      0.3203
! --------------------------------------
!  Double precision appears to have 16 digits of accuracy
!  Assuming 8 bytes per DOUBLEPRECISION word
! --------------------------------------
!  Timing calibration ; time =    49.60619788616896     hundredths of a second
!  Increase the size of the arrays if this is <30 
!   and your clock precision is =<1/100 second
!  ---------------------------------------------------
! Function     Rate (MB/s)  RMS time   Min time  Max time
! Assignment:   28.5786      0.1711      0.1680      0.1719
! Scaling   :   24.5776      0.2000      0.1953      0.2031
! Summing   :   23.9392      0.3035      0.3008      0.3125
! SAXPYing  :   23.0415      0.3180      0.3125      0.3281
! --------------------------------------
!  Double precision appears to have 16 digits of accuracy
!  Assuming 8 bytes per DOUBLEPRECISION word
! --------------------------------------
!  Timing calibration ; time =    49.99680034816265     hundredths of a second
!  Increase the size of the arrays if this is <30 
!   and your clock precision is =<1/100 second
!  ---------------------------------------------------
! Function     Rate (MB/s)  RMS time   Min time  Max time
! Assignment:   28.5786      0.1711      0.1680      0.1719
! Scaling   :   24.5776      0.1980      0.1953      0.1992
! Summing   :   23.9393      0.3057      0.3008      0.3398
! SAXPYing  :   23.0415      0.3185      0.3125      0.3437

---
		">NFS only behaves properly ...
		 ...when the computer is not drawing power."
			- Root Boy Jim
cks@hawkwind.utcs.toronto.edu	           ...!{utgpu,utzoo,watmath}!utgpu!cks

From tmaeno@cc.titech.ac.jp  Thu Oct  3 01:16:13 1991
Received: from rc.cc.titech.ac.jp ([131.112.4.59]) by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA03939; Thu, 3 Oct 91 01:16:13 EDT
From: tmaeno@cc.titech.ac.jp
Received: by rc.cc.titech.ac.jp (5.65+1.5W/r2TM)
	id AA17688; Thu, 3 Oct 91 14:16:38 JST
Return-Path: <tmaeno@cc.titech.ac.jp>
Message-Id: <9110030516.AA17688@rc.cc.titech.ac.jp>
Subject: memory bandwidth of ETA-10E and MIPS RC6280
To: mccalpin
Date: Thu, 3 Oct 91 14:16:38 GMT+9:00
X-Mailer: ELM [version 2.3 PL11]
Status: RO

Dear Prof. John D. McCalpin,

  I ran stream_{sd}.f on ETA-10E and MIPS RC6280.

  Very interesting.

			Toshinori Maeno
			Associate Professor
			Computer Center, Tokyo Institute of Technology
============================================================================

Date	1991-10-04
Run on ETA-10E (Computer Center, Tokyo Institute of Technology)
Clock speed 10.5ns
Using only one processor
ETA Unix
ftn77 103.7
vectorized

array size(n) 5000000

--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 8.0634412473998 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =
<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment: 2974.7025      0.0269      0.0269      0.0270
Scaling   : 3010.5016      0.0266      0.0266      0.0266
Summing   : 4357.4942      0.0276      0.0275      0.0276
SAXPYing  : 4214.7701      0.0285      0.0285      0.0285

=============================================================================

Run on MIPS RC6280 (Computer Center, Tokyo Institute of Technology)
  60 MHz
  16kB primary data cache
 512kB secondary cache

RiscOS/4.52
f77 2.11
  option -O -mips2
array size(n) = 400000
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    30.00000026077032     hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   53.3335      0.1232      0.1200      0.1400
Scaling   :   58.1815      0.1232      0.1100      0.1300
Summing   :   56.4706      0.1801      0.1700      0.1900
SAXPYing  :   56.4706      0.1801      0.1700      0.1900

From tarolli@westcoast.esd.sgi.com  Thu Oct  3 09:03:56 1991
Received: from SGI.COM by perelandra.cms.udel.edu (5.52/890607.SGI) (for mccalpin) id AA04350; Thu, 3 Oct 91 09:03:56 EDT
Received: from [192.26.51.77] by sgi.sgi.com via SMTP (910911.SGI.EXPERIMENTAL/910110.SGI) for mccalpin@perelandra.cms.udel.edu id AA20668; Thu, 3 Oct 91 06:02:30 -0700
Received: from westcoast.esd.sgi.com by mars.esd.sgi.com via SMTP (910711.SGI/910709.SGI.autocf) for @sgi.sgi.com:mccalpin@perelandra.cms.udel.edu id AA18322; Thu, 3 Oct 91 06:02:29 -0700
Received: by westcoast.esd.sgi.com (910524.SGI/900721.SGI) for @mars.esd.sgi.com:mccalpin@perelandra.cms.udel.edu id AA22570; Thu, 3 Oct 91 06:02:26 -0700
Date: Thu, 3 Oct 91 06:02:26 -0700
From: tarolli@westcoast.esd.sgi.com (Gary Tarolli)
Message-Id: <9110031302.AA22570@westcoast.esd.sgi.com>
To: mccalpin
Subject: Re: Attainable Memory Bandwidth (update)
In-Reply-To: your article <MCCALPIN.91Oct2094337@pereland.cms.udel.edu>
Status: RO

something in the numbers doesn't make sense.  The SGI Indigo should be
closer to the 4D/35.  The difference between the 2 machines is 33 vs 36Mhz
clock speed and 32K vs 64K cache.  The clock speed difference is only
about 10%.  If you are truly measuring memory thruput (and not cache
thruput), then the cache size should be largely irrelevant.  Unless
the 4D/35 has a data cache block size of 16 vs. 8 for the Indigo.

Are you sure your pgm copies enough data to render the cache size
irrelevant?



From alex@Think.COM  Thu Oct  3 09:52:13 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA04361; Thu, 3 Oct 91 09:52:13 EDT
Return-Path: <alex@Think.COM>
Received: from Berlin.Think.COM by mail.think.com; Thu, 3 Oct 91 09:50:46 -0400
Received: from Zeus.Think.COM by berlin.think.com; Thu, 3 Oct 91 09:50:44 -0400
From: Alex Vasilevsky <alex@Think.COM>
Received: by zeus.think.com; Thu, 3 Oct 91 09:50:42 EDT
Date: Thu, 3 Oct 91 09:50:42 EDT
Message-Id: <9110031350.AA16256@zeus.think.com>
To: mccalpin
Subject: CM2 data for stream code
Status: RO


Here is the CM-2 data for running the stream code.  The following switches
were used with CM Fortran compiler:

%cmf stream_d.fcm -o stream

The global optimizer was disabled due to the way the code was written, the
global optimizer will remove all of this code out of the loop since it is
loop invarient and then the timings will be totally bogus.

The software used for this was the following:

1. CM Fortran Compiler Release 1.1 
2. CMost Release 6.1 Beta II

The machine used was CM2 64K, CM2 32K and CM2 16K all running at 7Mhz.

	-Alex V.

P.S.  Here is all the data.


xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
CM2 64K
-------------------------------------
Double precision appears to have 16 digits of accuracy
Assuming 8 bytes per DOUBLEPRECISION word
-------------------------------------
Calibrating CM timer...Done. CM speed = 7.00 MHz
Timing calibration ; time =   143.324911594391 hundredths of a second
Increase the size of the arrays if this is <30  and your clock precision is =<1
/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment  56555.9654       .0231       .0231       .0231
Scaling     56555.9654       .0231       .0231       .0231
Summing     81244.0132       .0242       .0242       .0242
SAXPYing    80593.2540       .0244       .0244       .0244
xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
CM2 32K
-------------------------------------
Double precision appears to have 16 digits of accuracy
Assuming 8 bytes per DOUBLEPRECISION word
-------------------------------------
Calibrating CM timer...Done. CM speed = 7.00 MHz
Timing calibration ; time =   143.324911594391 hundredths of a second
Increase the size of the arrays if this is <30  and your clock precision is =<1
/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment  28277.9821       .0231       .0231       .0231
Scaling     28277.8832       .0231       .0231       .0231
Summing     40622.0066       .0242       .0242       .0242
SAXPYing    40296.6270       .0244       .0244       .0244
xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
CM2 16K
-------------------------------------
Double precision appears to have 16 digits of accuracy
Assuming 8 bytes per DOUBLEPRECISION word
-------------------------------------
Calibrating CM timer...Done. CM speed = 7.00 MHz
Timing calibration ; time =   143.325054645538 hundredths of a second
Increase the size of the arrays if this is <30  and your clock precision is =<1
/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment  14138.9911       .0231       .0231       .0231
Scaling     14138.9402       .0231       .0231       .0231
Summing     20311.0033       .0242       .0242       .0242
SAXPYing    20148.3135       .0244       .0244       .0244
xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx




From rtp1@quads.uchicago.edu Thu Oct  3 11:47:37 1991
From: rtp1@quads.uchicago.edu (raymond thomas pierrehumbert)
Newsgroups: comp.benchmarks
Subject: Re: Attainable Memory Bandwidth (update)
Message-Id: <9110031350.BB16256@quads.uchicago.edu>
Date: 3 Oct 91 02:44:04 GMT
Organization: University of Chicago
Status: RO


For fans of antique hardware, I brought my aHPollo DN10k out of moth-
balls long enough to run the stream benchmark.  The results are
really interesting (followup to alt.folklore.computers).  I include
the DN10k results below.

                              Actual Transfer Rate (MB/s)
Machine                   Copy     SSCAL       Sum     SAXPY
------------------------------------------------------------
Cray Y/MP 8 cpu        19291.6   19294.2   26588.9   26802.2
Cray Y/MP 4 cpu         9685.8    9678.9   13781.4   13851.2
Cray Y/MP               2426.4    2426.2    3454.4    3396.9
IBM RS6000-950           193.9     177.8     195.9     192.0
IBM RS6000-530           114.3     106.7     109.1     114.3
IBM RS6000-320            61.5      61.5      60.0      60.0
SGI 4D/35                 53.5      36.9      48.0      42.4
HP 9000/730               53.3      48.0      55.4      55.4
HP 9000/720               43.6      40.0      45.0      42.3
aHPollo DN10010       41.        48.        54.        54.
SGI Indigo                34.3      22.9      27.7      26.7
Sun 670                   28.2      32.0      30.0      30.0
DEC 5000                  25.6      24.4      24.0      22.4
Sun 4/490                 25.0      25.8      25.5      24.5
SGI 4D/240                18.4      16.5      19.8      19.0
Omron Luna88k             14.4      14.4      16.0      13.1
Sun SS1                   14.1      12.3      13.6      13.1
SGI 4D/25                 12.7      10.1      10.4      10.0
------------------------------------------------------------

Note the position of the DN10k relative to the "hot" new HP
machines.  The Apollo architecture, which HP bought and then
allowed to die (for a variety of reasons) actually matches or
beats the snakes.  This is on a machine first released three years
before the snakes.  The 2x processor, if it had not been cancelled
for technical reasons, would have blown away the snakes on
bandwidth-limited applications, and been competitive with the IBM
machines.

This confirms my opinion that HP should have worked much harder to
keep the PRISM architecture alive.

From alex@Think.COM  Thu Oct  3 12:29:33 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA04715; Thu, 3 Oct 91 12:29:33 EDT
Return-Path: <alex@Think.COM>
Received: from Berlin.Think.COM by mail.think.com; Thu, 3 Oct 91 12:28:10 -0400
Received: from Zeus.Think.COM by berlin.think.com; Thu, 3 Oct 91 12:28:09 -0400
From: Alex Vasilevsky <alex@Think.COM>
Received: by zeus.think.com; Thu, 3 Oct 91 12:28:07 EDT
Date: Thu, 3 Oct 91 12:28:07 EDT
Message-Id: <9110031628.AA16372@zeus.think.com>
To: mccalpin
In-Reply-To: John D. McCalpin's message of Thu, 3 Oct 91 11:40:08 EDT <9110031540.AA04563@perelandra.cms.udel.edu>
Subject:  CM2 data for stream code
Status: RO

   Date: Thu, 3 Oct 91 11:40:08 EDT
   From: mccalpin@perelandra.cms.udel.edu (John D. McCalpin)

   Thanks a lot!  This is great stuff!   
   The most humorous column in the spreadsheet is the one 
   that measures memory throughput in bytes/clock cycle.
   Most of the "Killer Micros" are around 1 byte/cycle, with
   the Y/MP-8 up to 160 bytes/cycle.  The 64k CM2 does over
   11,500 bytes/cycle!!!!
   --
   John D. McCalpin			mccalpin@perelandra.cms.udel.edu
   Assistant Professor			mccalpin@brahms.udel.edu
   College of Marine Studies, U. Del.	DELOCN::MCCALPIN (SPAN)


You welcome.  Actually, there is a CM200 machine around, that runs at 10
Mhz which should give better memory bandwith performance, but I did not
have an access to it last night to run the code on it.  

I really like your program though, I wish we had it available last winter
when I was working on Gordon Bell award and was trying to measure Sun cache
performance.


	-Alex V.

From tarolli@tenno.boston.sgi.com  Thu Oct  3 12:32:21 1991
Received: from SGI.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA04731; Thu, 3 Oct 91 12:32:21 EDT
Received: from relay.sgi.com by sgi.sgi.com via SMTP (910911.SGI.EXPERIMENTAL/910110.SGI)
	for mccalpin@perelandra.cms.udel.edu id AA28184; Thu, 3 Oct 91 09:30:56 -0700
Received: from relay.boston.sgi.com by relay.sgi.com (5.52/900423.SGI)
	for @sgi.sgi.com:mccalpin@perelandra.cms.udel.edu id AA07519; Thu, 3 Oct 91 09:30:53 PDT
Received: from tenno.boston.sgi.com by sgibos.boston.sgi.com (5.52/900721.SGI)
	for @relay.sgi.com:mccalpin@perelandra.cms.udel.edu id AA04565; Thu, 3 Oct 91 12:30:22 EDT
Received: by tenno.boston.sgi.com (910711.SGI/900721.SGI)
	for @sgibos.boston.sgi.com:mccalpin@perelandra.cms.udel.edu id AA09560; Thu, 3 Oct 91 12:30:17 -0400
Date: Thu, 3 Oct 91 12:30:17 -0400
From: tarolli@tenno.boston.sgi.com (Gary Tarolli)
Message-Id: <9110031630.AA09560@tenno.boston.sgi.com>
To: olson@esd.sgi.com, mccalpin
Subject: Re: Attainable memory bandwidth
Status: RO

I reran the tests from Dave Olson's machines and got the following results
on my Indigo:

the following results were obtained by compiling on my PI running 3.3:

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    52.00000219047070     hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   36.9233      0.1321      0.1300      0.1400
Scaling   :   34.2857      0.1481      0.1400      0.1600
Summing   :   36.0000      0.2051      0.2000      0.2100
SAXPYing  :   36.0000      0.2111      0.2000      0.2200

the following results were obtained by compiling on my Indigo running 4.0

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    51.99999623000622     hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   36.9231      0.1310      0.1300      0.1400
Scaling   :   26.6667      0.1901      0.1800      0.2000
Summing   :   32.7273      0.2230      0.2200      0.2300
SAXPYing  :   30.0000      0.2480      0.2400      0.2500

Obviously there's a bug in the 4.0 compiler that produces suboptimal code.
However, even given that, I got better numbers than Dave.  I am going to
submit a bug reporting that the performance of the code produced by the 4.0
compiler went down hill.  I have look at the executable, and the 3.3 simply
does better floating point instruction scheduling.
______________________________________________________________________________
  _____              ______           _  _	(508)562-4800  tarolli@sgi.com
 / ___  __  __         / __  __  ___ // // *	M/S DER-200
(____/ (_/_/ (_(_/    / (_/_/ (_(_/ (/_(/_/_
	       _/




From ECF_STBO@jhuvms.hcf.jhu.edu  Thu Oct  3 15:38:36 1991
Received: from jhuvms.hcf.jhu.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA05214; Thu, 3 Oct 91 15:38:36 EDT
Date: Thu, 3 Oct 91 15:38 EDT
From: "Save the Bay: shoot a developer" <ECF_STBO@jhuvms.hcf.jhu.edu>
To: mccalpin
Message-Id: <06E04A2211DF00C86B@JHUVMS.BITNET>
X-Envelope-To: mccalpin@perelandra.cms.udel.edu
X-Vms-To: IN%"mccalpin@perelandra.cms.udel.edu"
Status: RO

I have run the memory bandwidth code on a VAX 6000-410 equipped with a vector
processor running VMS 5.4-2. The memory bus has a theoretical peak of something
like 100MB/s. Unfortunately, there are usually a bunch of other users
on the machine so the timer I'm using (which measures wall time) varies a lot.
I tried to run the test in the morning when the system isn't very busy.
I had to make some changes in the code. I had to put a call to a dummy
subroutine after each timed loop, otherwise the compiler would optimize them
out. The timer routine was changed to use the secnds() routine rather than
etime. I also changed the size of the problem to 100000 because I didn't want
to hog too much memory on the system. The resulting working set size was a
couple of MB. 

-----


compiled without using vector processor:
-------------------------------------
Double precision appears to have 17 digits of accuracy
Assuming 8 bytes per DOUBLEPRECISION word
-------------------------------------
Timing calibration ; time =    131.2500000000000     hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
unction     Rate (MB/s)  RMS time   Min time  Max time
ssignment:   12.4121      0.3580      0.1289      0.7383
caling   :    8.9043      0.3653      0.1797      0.4922
umming   :    9.1701      0.5277      0.2617      0.7578
AXPYing  :    7.9792      0.5586      0.3008      0.8086


compiled using vector processor:
-------------------------------------
Double precision appears to have 17 digits of accuracy
Assuming 8 bytes per DOUBLEPRECISION word
-------------------------------------
Timing calibration ; time =    60.93750000000000     hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
unction     Rate (MB/s)  RMS time   Min time  Max time
ssignment:   58.5143      0.0375      0.0273      0.0703
caling   :   58.5143      0.0485      0.0273      0.1016
umming   :   61.4400      0.0669      0.0391      0.0977
AXPYing  :   51.2000      0.0660      0.0469      0.0977

From csrcb@shark.mel.dit.csiro.au  Thu Oct  3 19:12:33 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA06603; Thu, 3 Oct 91 19:12:33 EDT
Received: from trout.mel.dit.CSIRO.AU by shark.mel.dit.csiro.au with SMTP id AA01933
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Fri, 4 Oct 91 09:11:09 +1000
Received: by trout.mel.dit.CSIRO.AU (4.1/SMI-4.0)
	id AA09958; Fri, 4 Oct 91 09:11:08 EST
Date: Fri, 4 Oct 91 09:11:08 EST
From: Robert.Bell@mel.dit.csiro.au
Message-Id: <9110032311.AA09958@trout.mel.dit.CSIRO.AU>
To: cmg@cray.com, mccalpin
Subject: Benchmarks
Status: RO

John,
     Thanks for keeping the issue alive, and replying to Mark Boolootian.
     In your next post, it might be worth clarifying for Mark who seems
 interested in i/o, that each Cray CPU has a fourth port to memory for i/o.
 I've no idea how to test that in parallel with a SAXPY.
       Charles Grassl has provided me with results from the C-90.
     You are free to post them, and it would be good to do so, but point
 out that the memory is still far short of the final configuration.
     Note that they are up to 7 CPUs.
     Cheers,
            Rob. Bell.

From shair@ux2.cso.uiuc.edu  Thu Oct  3 20:44:48 1991
Received: from ux2.cso.uiuc.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA06618; Thu, 3 Oct 91 20:44:48 EDT
Received: by ux2.cso.uiuc.edu id AA20184
  (5.65c/IDA-1.4.4 for mccalpin@perelandra.cms.udel.edu); Thu, 3 Oct 1991 19:43:27 -0500
Date: Thu, 3 Oct 1991 19:43:27 -0500
From: Robert Shair - IBM <shair@ux2.cso.uiuc.edu>
Message-Id: <199110040043.AA20184@ux2.cso.uiuc.edu>
To: mccalpin
Subject: Re: Attainable memory bandwidth
Newsgroups: comp.benchmarks
References: <MCCALPIN.91Sep27212326@pereland.cms.udel.edu> <MCCALPIN.91Oct2094028@pereland.cms.udel.edu>
Status: RO

Trying your bandwidth test on RISC6000 540 and 320H.  Might also run on
a Sequent, if I can ever find it quiet.

What optimizations have you used on the RISC systems?  I'd better do at 
least as well.

Am running xlf V2.1 with no maintenance applied.


From csrcb@shark.mel.dit.csiro.au  Fri Oct  4 03:07:03 1991
Received: from shark.mel.dit.CSIRO.AU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA06718; Fri, 4 Oct 91 03:07:03 EDT
Received: by shark.mel.dit.csiro.au id AA06405
  (5.65b/IDA-1.4.3/DIT-1.2 for mccalpin@perelandra.cms.udel.edu); Fri, 4 Oct 91 17:05:41 +1000
Date: Fri, 4 Oct 91 17:05:41 +1000
From: Robert Bell <Robert.Bell@mel.dit.csiro.au>
Message-Id: <9110040705.AA06405@shark.mel.dit.csiro.au>
To: mccalpin
Subject: Memory benchmarks
Status: RO

 John,
      I have refined my benchmark code, and put the code and sample outputs
 for access by anonymous ftp on shark.mel.dit.csiro.au, in directory
 staff/csrcb/bm20only.  The present version is for a Cray.  The code in bm1.f
 needs to be changed for different machines.  bm2.f just tests the bm1.f
 routines, and bm20.f does the memory benchmarks.   Sample outputs and summaries
 are available.
      Charles Grassl from Cray is likely to provide C-90 results.
      An autotasking run output is available, but not for a dedicated machine.
      Regards,
              Rob. Bell.

From jacobsd@ucs.orst.edu  Fri Oct  4 07:21:37 1991
Received: from ucs.orst.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA07083; Fri, 4 Oct 91 07:21:37 EDT
Received: from ucs.orst.edu by beasley.ucs.orst.edu (4.1/SMI-3.2)
	id AA29166; Fri, 4 Oct 91 04:20:16 PDT
Received: by ucs.orst.edu (4.1/SMI-4.1)
	id AA27262; Fri, 4 Oct 91 04:14:39 PDT
Date: Fri, 4 Oct 91 04:14:39 PDT
From: jacobsd@ucs.orst.edu (Dana Jacobsen)
Message-Id: <9110041114.AA27262@ucs.orst.edu>
To: mccalpin
Subject: stream_d.f timing for FPS 511
Cc: jacobsd@cs.orst.edu
Status: RO


  This is the timing for stream_d.f, run on an FPS 511.  The FPS is configured
with 1 SPARC scalar processor (of possible 8), 1 vector processor (of possible 2),
and 0 matrix processors (of possible 168 (i860s)).  128 Meg of RAM.  Running
+FPX 5.0.1.   I changed n to 900000 to get better numbers.  I do not know
all the tricks of this compiler, so it is possible I've missed some key compiler
flags that would speed this up -- I just told it to make vectorized code.
  Unfortunately it looks like FPS is going to go out of business.  Sigh.

========
Script started on Fri Oct  4 04:05:48 1991
mesg: cannot change mode
/dev/ttyp1: Not owner
fps /home/ucs/u1/staff/jacobsd/src/bench/mem 401% ls
stream.f        stream_d.f      stream_d.o      stream_s.f      table.print     table.ps        table.sc        typescript
fps /home/ucs/u1/staff/jacobsd/src/bench/mem 402% f77 -Oc vec+ stream_d.f
stream_d.f:
   MAIN stream:
   second:
   realsize:
   dummy:
7.3u 1.5s 0:15 56% 0+1776k 12+90io 0pf+0w
fps /home/ucs/u1/staff/jacobsd/src/bench/mem 403% ./a.out
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     126.00000537932 hundredths of a second
Increase the size of the arrays if this is <30
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  288.0016      0.0591      0.0500      0.0600
Scaling   :  180.0002      0.0871      0.0800      0.0900
Summing   :  216.0002      0.1122      0.1000      0.1200
SAXPYing  :  196.3642      0.1121      0.1100      0.1200
4.8u 0.2s 0:11 46% 0+18016k 0+0io 0pf+0w
fps /home/ucs/u1/staff/jacobsd/src/bench/mem 404% head -50 stream_d.f
* Program: Stream
* Programmer: John D. McCalpin
* Revision: 2.0, September 30,1991
*
* This program measures memory transfer rates in MB/s for simple
* computational kernels coded in Fortran.  These numbers reveal the
* quality of code generation for simple uncacheable kernels as well
* as showing the cost of floating-point operations relative to memory
* accesses.
*
* INSTRUCTIONS:
*       1) Stream requires a cpu timing function called second().
*          A sample is shown below.  This is unfortunately rather
*          system dependent.  It helps to know the granularity of the
*          timing.  The code below assumes that the granularity is
*          1/100 seconds.
*       2) Stream requires a good bit of memory to run.
*          Adjust the Parameter 'N' in the second line of the main
*          program to give a 'timing calibration' of at least 20 clicks.
*          This will provide rate estimates that should be good to
*          about 5% precision.
*       3) Compile the code with full optimization.  Many compilers
*          generate unreasonably bad code before the optimizer tightens
*          things up.  If the results are unreasonable good, on the
*          other hand, the optimizer might be too smart for me!
*       4) Mail the results to mccalpin@perelandra.cms.udel.edu
*          Be sure to include:
*               a) computer hardware model number and software revision
*               b) the compiler flags
*               c) all of the output from the test case.
*
* Thanks!
*
      PROGRAM stream
C     .. Parameters ..
      INTEGER n,ntimes
      PARAMETER (n=900000,ntimes=10)
C     ..
C     .. Local Scalars ..
      DOUBLE PRECISION t,t0
      INTEGER j,k,nbpw
C     ..
C     .. Local Arrays ..
      DOUBLE PRECISION a(n),b(n),c(n),maxtime(4),mintime(4),rmstime(4),
     $                 times(4,ntimes)
      INTEGER bytes(4)
      CHARACTER label(4)*11
C     ..
C     .. External Functions ..
      DOUBLE PRECISION second
fps /home/ucs/u1/staff/jacobsd/src/bench/mem 405% exit
fps /home/ucs/u1/staff/jacobsd/src/bench/mem 406%
script done on Fri Oct  4 04:07:00 1991
========
--
Dana Jacobsen
jacobsd@cs.orst.edu
Oregon State University     Computer Science

From uunet.UU.NET!ssi!rousay!dgstren  Fri Oct  4 10:48:16 1991
Received: from relay2.UU.NET by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA07564; Fri, 4 Oct 91 10:48:16 EDT
Received: from uunet.uu.net (via LOCALHOST.UU.NET) by relay2.UU.NET with SMTP 
	(5.61/UUNET-internet-primary) id AA09685; Fri, 4 Oct 91 10:47:03 -0400
Received: from ssi.UUCP by uunet.uu.net with UUCP/RMAIL
	(queueing-rmail) id 104505.9605; Fri, 4 Oct 1991 10:45:05 EDT
Received: from rousay.com by ssi (4.0/SMI-4.0)
	id AA28159; Fri, 4 Oct 91 09:44:15 CDT
Received: by rousay.com (4.0/SMI-4.0)
	id AA23255; Fri, 4 Oct 91 09:44:12 CDT
Date: Fri, 4 Oct 91 09:44:12 CDT
From: uunet.UU.NET!ssi!rousay!dgstren (Dave Strenski)
Message-Id: <9110041444.AA23255@rousay.com>
To: uunet.UU.NET!uunet!perelandra.cms.udel.edu!mccalpin
Subject: Memory Band width
Status: RO

John D. McCalpin,
	I am unable to do anonymous ftp's, could you please send/mail me the
	following files from your memory band width tests.
		bench/stream/stream_d.f
		bench/pub/stream/Results/*
		
							Thank You, 
							Dave Strenski

From jacobsd@frisby.CS.ORST.EDU  Fri Oct  4 10:51:06 1991
Received: from frisby.CS.ORST.EDU by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA07574; Fri, 4 Oct 91 10:51:06 EDT
Received: from frisby.CS.ORST.EDU by frisby.CS.ORST.EDU (4.1/1.30)
	id AA12234; Fri, 4 Oct 91 07:49:40 PDT
Message-Id: <9110041449.AA12234@frisby.CS.ORST.EDU>
To: "John D. McCalpin" <mccalpin>
Cc: Dana Jacobsen <jacobsd@ucs.orst.edu>
Subject: Re: stream_d.f timing for FPS 511 
In-Reply-To: Your message of Fri, 04 Oct 91 08:51:20 -0400.
             <9110041251.AA07180@perelandra.cms.udel.edu> 
Date: Fri, 04 Oct 91 07:49:36 PDT
From: jacobsd@frisby.CS.ORST.EDU
Status: RO

> Do you happen to know the clock speed used on this box?
> One of the columns in my summary table is bandwidth in bytes/clock cycle.

  HZ is defined to be 60 in /usr/include/sys/param.h.

  The FPS literature states a "15ns clock" and 67 (native) MIPS.

  The vector processor is rated at 67 MFLOPS (judging from the other figures
and it's linpack performance, these are peak MFLOPS)
  8 vector registers, 1024 64-bit elements.

 more propaganda: 
   Scalable Interconnect Architecture:
     33MHz cycle time, 1GB/sec bandwidth
   Memory:
     Multiple 267 MB/sec peak transfer paths


  If you'd like any more information, I can ask some of the technical support
people.

Disclaimer:  I don't work for FPS -- my university bought a system, and I get
to play with it.  I'd rather another Oregon company not go under however..
Opinion:  It's fast but quirky (compared to a normal Sun box).  
--
Dana Jacobsen                      Oregon State University
jacobsd@cs.orst.edu                  Computer Science
..!hplabs!hp-pcd!orstcs!jacobsd
              Stop corruption in government now: Abolish the Republican Party!

From pmk@craycos.com  Fri Oct  4 14:57:31 1991
Received: from aspen.craycos.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA08765; Fri, 4 Oct 91 14:57:31 EDT
Received: from copper.craycos.com by aspen.Craycos.COM (4.1/TotalHack-4.0)
	id AA25558; Fri, 4 Oct 91 12:55:30 MDT
Received: from zymurgy.Craycos.COM by copper.craycos.com (4.0/SMI-4.0)
	id AA21746; Fri, 4 Oct 91 12:55:17 MDT
Date: Fri, 4 Oct 91 12:55:17 MDT
From: pmk@craycos.com (Peter Klausler)
Message-Id: <9110041855.AA21746@copper.craycos.com>
Received: by zymurgy.Craycos.COM (4.0/SMI-4.0)
	id AA04961; Fri, 4 Oct 91 12:55:14 MDT
To: mccalpin
Subject: Re: Attainable memory bandwidth
Newsgroups: comp.arch,comp.benchmarks
In-Reply-To: <MCCALPIN.91Oct3160250@pereland.cms.udel.edu>
References: <MCCALPIN.91Sep27212326@pereland.cms.udel.edu> <MCCALPIN.91Oct2094028@pereland.cms.udel.edu> <108498@lll-winken.LLNL.GOV>
Organization: Cray Computer Corporation
Cc: 
Status: RO

In article <MCCALPIN.91Oct3160250@pereland.cms.udel.edu> you write:
>Each cpu of the Y/MP is capable of two 64-bit loads and one 64-bit
>store per clock cycle --- these are overlappable vector operations.
>Thus each cpu can move 24 bytes/cycle, or 4000 MB/s.

There are four memory ports of an X-MP/Y-MP, and their functions are:

	port A:	B-register block loads, vector loads (2nd choice), exchange load
	port B: T-register block loads, vector loads (1rst choice)
	port C: Scalar loads, scalar/B/T/vector stores, exchange store
	port D: instruction fetch, all I/O

Vector loads will use port B if free, else port A. All stores, except from
I/O channels, go out through port C. (On an X-MP instruction fetch uses port A.)

So you need to include I/O bandwidth in your estimate, as all Cray machines
(1/X/Y/2/3) perform I/O through processor memory ports.

An important feature of CRAY-1/X/Y memory ports is that references run
serially. If, say, a vector element load is waiting for a busy section or bank,
its port will wait for the element to proceed. A port won't go on ahead to
process other element references that could perhaps proceed. This serial
operation is important for the implementation of vector chaining operations,
since the data is guaranteed to flow back to the processor from a vector load
or gather in the proper order, possibly with gaps.

The CRAY-2 and CRAY-3 memory systems are very different from those of the
CRAY-1 and its descendents. The CRAY-2 considers memory to be divided in four
quadrants and has one quarter-speed port per processor for each quadrant. All
memory references are single-element loads and stores, and the references
constituting a block load or store are able to get out of order.

Peter Klausler
Compilers
Cray Computer (NOT Cray Research)
pmk@craycos.com

From UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper  Fri Oct  4 14:55:34 1991
Received: from ucsd.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA08761; Fri, 4 Oct 91 14:55:34 EDT
Received: from celit.UUCP by ucsd.edu; id AA10425
	sendmail 5.64/UCSD-2.2-sun via UUCP
	Fri, 4 Oct 91 10:53:12 -0700
Received: by celit.fps.com (5.51/celerity1.1)
	id AA07498; Fri, 4 Oct 91 10:25:04 PDT for mccalpin@perelandra.cms.udel.edu at ucsd
Posted-Date: Fri, 4 Oct 91 10:22:56 PDT
Received: from keeper.fps_net by fpssun (4.1/SMI-4.1)
	id AA10317; Fri, 4 Oct 91 10:16:27 PDT
Received: by keeper.fps_net (4.1/SMI-4.1)
	id AA09181; Fri, 4 Oct 91 10:22:56 PDT
Date: Fri, 4 Oct 91 10:22:56 PDT
From: UCSD.EDU!celit!fpssun.fps.com!keeper.fps.COM!keeper (Brian Whitney)
Message-Id: <9110041722.AA09181@keeper.fps_net>
To: mccalpin
Subject: Megabytes results from FPS
Status: RO


Here are numerous results I have obtained from machines
available to me here at FPS.  I am a bit at a loss as to
why my SS1+ times are so much different compared to yours.

Brian Whitney
FPS Computing

keeper@fps.com

##################################
470/real

This test was run on a SUN 470 running 
SunOS 4.1.1 with 65Mbytes of main memory.

--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     56.0000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   17.3913      0.2300      0.2300      0.2300
Scaling   :   13.7931      0.2910      0.2900      0.3000
Summing   :   16.2162      0.3760      0.3700      0.3800
SAXPYing  :   18.1818      0.3390      0.3300      0.3400
**********************************
##################################
470/double

This test was run on a SUN 470 running 
SunOS 4.1.1 with 65Mbytes of main memory.

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     65.000002831221 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   21.8182      0.2300      0.2200      0.2400
Scaling   :   24.0000      0.2090      0.2000      0.2100
Summing   :   20.5714      0.3570      0.3500      0.3600
SAXPYing  :   20.0000      0.3660      0.3600      0.3700
**********************************
##################################
sps1/real

This test was run on a SUN SPARCStation 1+ running
SunOS 4.1.

hw1% !!
sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     149.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   12.9032      0.3120      0.3100      0.3200
Scaling   :   11.4286      0.3550      0.3500      0.3600
Summing   :   12.7660      0.4811      0.4700      0.5000
SAXPYing  :   12.5000      0.4850      0.4800      0.4900
**********************************
##################################
sps1/double

This test was run on a SUN SPARCStation 1+ running
SunOS 4.1.

hw1% !!
dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     181.00000321865 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   16.0000      0.3141      0.3000      0.3300
Scaling   :   16.0000      0.3100      0.3000      0.3200
Summing   :   14.4000      0.5061      0.5000      0.5200
SAXPYing  :   13.3333      0.5540      0.5400      0.5600
**********************************
##################################
ipc/real

This test was run on a SUN SPARCStation IPC running
SunOS 4.1.1 with 12Mbytes of main memory.

keeper% !!
sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     161.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   12.9032      0.3140      0.3100      0.3200
Scaling   :   11.4286      0.3540      0.3500      0.3600
Summing   :   12.5000      0.4800      0.4800      0.4800
SAXPYing  :   12.5000      0.4830      0.4800      0.4900
**********************************
##################################
ipc/double

This test was run on a SUN SPARCStation IPC running
SunOS 4.1.1 with 12Mbytes of main memory.

keeper% !!
dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     194.00001168251 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   15.4839      0.3296      0.3100      0.4500
Scaling   :   15.4839      0.3151      0.3100      0.3400
Summing   :   14.4000      0.5093      0.5000      0.5600
SAXPYing  :   13.0909      0.5552      0.5500      0.6000
**********************************
##################################
ipc-48/real

This test was run on a SUN SPARCStation IPC running
SunOS 4.1.1 with 48Mbytes of main memory.

rcl% !!
sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     156.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   12.9032      0.3160      0.3100      0.3200
Scaling   :   11.7647      0.3541      0.3400      0.3600
Summing   :   12.5000      0.4840      0.4800      0.4900
SAXPYing  :   12.5000      0.4880      0.4800      0.4900
**********************************
##################################
ipc-48/double

This test was run on a SUN SPARCStation IPC running
SunOS 4.1.1 with 48Mbytes of main memory.

rcl% !!
dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     186.99999749660 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   15.4839      0.3180      0.3100      0.3200
Scaling   :   15.4839      0.3140      0.3100      0.3200
Summing   :   14.4000      0.5050      0.5000      0.5100
SAXPYing  :   13.0909      0.5520      0.5500      0.5600
**********************************
##################################
510ea/real

This test was run on an FPS Model 510 EA running FPX 4.3.3
with 256Mbytes of main memory.  This is a scalar only machine,
no vector processor running at 33Mhz.

thor% !!
sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
  Timing calibration ; time =    65.0483   hundredths   of a second
  Increase the size of the arrays if this is <30 
   and your clock precision is =<1/100 second
  ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   14.0845      0.2876      0.2840      0.2970
Scaling   :   11.5607      0.3496      0.3460      0.3575
Summing   :   13.0592      0.4635      0.4594      0.4695
SAXPYing  :   11.8460      0.5101      0.5065      0.5135
**********************************
##################################
510ea/double

This test was run on an FPS Model 510 EA running FPX 4.3.3
with 256Mbytes of main memory.  This is a scalar only machine,
no vector processor running at 33Mhz.

thor% !!
dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
  Timing calibration ; time =    59.750800000000   hundredths   of a second
  Increase the size of the arrays if this is <30 
   and your clock precision is =<1/100 second
  ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   26.0874      0.1865      0.1840      0.1919
Scaling   :   21.6216      0.2244      0.2220      0.2294
Summing   :   23.6842      0.3090      0.3040      0.3245
SAXPYing  :   21.3739      0.3387      0.3369      0.3407
**********************************
##################################
511ea/real

This test was run on an FPS Model 511 EA running FPX 4.3.3
with 256Mbytes of main memory.  This machine includes a 
vector processor and runs at 33Mhz.

thor% !!
sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
  Timing calibration ; time =    45.0474   hundredths   of a second
  Increase the size of the arrays if this is <30 
   and your clock precision is =<1/100 second
  ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  121.2122      0.0336      0.0330      0.0355
Scaling   :   83.3334      0.0482      0.0480      0.0485
Summing   :   93.0247      0.0646      0.0645      0.0650
SAXPYing  :   93.0233      0.0651      0.0645      0.0675
**********************************
##################################
511ea/double

This test was run on an FPS Model 511 EA running FPX 4.3.3
with 256Mbytes of main memory.  This machine includes a 
vector processor and runs at 33Mhz.

thor% !dx
dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
  Timing calibration ; time =    49.550100000000   hundredths   of a second
  Increase the size of the arrays if this is <30 
   and your clock precision is =<1/100 second
  ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  246.1791      0.0199      0.0195      0.0275
Scaling   :  168.4270      0.0289      0.0285      0.0335
Summing   :  187.0178      0.0390      0.0385      0.0410
SAXPYing  :  187.0276      0.0391      0.0385      0.0420
**********************************
##################################
510s/real

This test was run on an FPS Model 510 SPARC running FPX+ 5.0.1
with 256Mbytes of main memory.  This machine is based on a
66.7 Mhz SPARC chip set from BIT.

--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     59.3890 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   37.2162      0.1196      0.1075      0.1390
Scaling   :   31.0382      0.1411      0.1289      0.1630
Summing   :   27.4907      0.2363      0.2183      0.2743
SAXPYing  :   26.5874      0.2476      0.2257      0.2730
**********************************
##################################
510s/double

This test was run on an FPS Model 510 SPARC running FPX+ 5.0.1
with 256Mbytes of main memory.  This machine is based on a
66.7 Mhz SPARC chip set from BIT.

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     76.876001060009 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   47.9176      0.1017      0.1002      0.1040
Scaling   :   44.9830      0.1085      0.1067      0.1119
Summing   :   36.4802      0.2011      0.1974      0.2081
SAXPYing  :   34.2300      0.2131      0.2103      0.2186
**********************************
##################################
510s-sun/real

This test was run on an FPS Model 510 SPARC running FPX+ 5.0.1
with 256Mbytes of main memory.  This machine is based on a
66.7 Mhz SPARC chip set from BIT.

This test was run using the SUN executable used in the previous
SUN SPARC tests.

--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     56.9606 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   30.0629      0.1352      0.1331      0.1391
Scaling   :   22.3537      0.1809      0.1789      0.1842
Summing   :   24.8278      0.2452      0.2417      0.2512
SAXPYing  :   25.9581      0.2338      0.2311      0.2382
**********************************
##################################
510s-sun/double

This test was run on an FPS Model 510 SPARC running FPX+ 5.0.1
with 256Mbytes of main memory.  This machine is based on a
66.7 Mhz SPARC chip set from BIT.

This test was run using the SUN executable used in the previous
SUN SPARC tests.

--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     67.112995684147 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   37.4373      0.1286      0.1282      0.1292
Scaling   :   36.2188      0.1329      0.1325      0.1332
Summing   :   27.9998      0.2577      0.2571      0.2582
SAXPYing  :   27.2899      0.2645      0.2638      0.2653
**********************************
##################################
511s/real

This test was run on an FPS Model 511 SPARC running FPX+ 5.0.1
with 256Mbytes of main memory.  This machine is based on a
66.7 Mhz SPARC chip set from BIT and includes a vector 
processor.

systst2% sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     36.8260 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  124.6492      0.0324      0.0321      0.0333
Scaling   :   85.3277      0.0471      0.0469      0.0475
Summing   :   95.0346      0.0635      0.0631      0.0640
SAXPYing  :   95.0360      0.0635      0.0631      0.0654
**********************************
##################################
511s/double

This test was run on an FPS Model 511 SPARC running FPX+ 5.0.1
with 256Mbytes of main memory.  This machine is based on a
66.7 Mhz SPARC chip set from BIT and includes a vector 
processor.

systst2% dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     58.325901627541 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  250.4950      0.0192      0.0192      0.0193
Scaling   :  170.8307      0.0282      0.0281      0.0283
Summing   :  190.5371      0.0381      0.0378      0.0386
SAXPYing  :  190.2197      0.0379      0.0379      0.0381
**********************************
##################################
mcp101/real

This test was run on an FPS MCP104 running release 2.1
with 64Mbytes of matrix memory.  This machine contains
4 Intel i860 running at 40Mhz on 1 bus.  The test was
run on a single i860.

systst2% !!
mrun sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time =    7.812500     hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  154.5660      0.0261      0.0259      0.0264
Scaling   :  103.6962      0.0386      0.0386      0.0391
Summing   :  115.9245      0.0522      0.0518      0.0522
SAXPYing  :   93.0909      0.0646      0.0645      0.0649
**********************************
##################################
mcp101/double

This test was run on an FPS MCP104 running release 2.1
with 64Mbytes of matrix memory.  This machine contains
4 Intel i860 running at 40Mhz on 1 bus.  The test was
run on a single i860.

warp% mrun dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    6.291955000000371      hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  154.1322      0.0312      0.0311      0.0315
Scaling   :  153.9616      0.0312      0.0312      0.0312
Summing   :  131.1701      0.0549      0.0549      0.0549
SAXPYing  :  130.8698      0.0550      0.0550      0.0550
**********************************
##################################
mcp707/real

This test was run on an FPS MCP728 running release 2.1
with 64Mbytes of matrix memory.  This machine contains
28 Intel i860 running at 40Mhz on 7 busses.  The test was
run on 7 i860s, each with its own bus.

setenv MCPCONFIG 7,1
systst2% !mr
mrun sxfer
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time =    9.018345000004047      hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment: 1046.1756      0.0038      0.0038      0.0038
Scaling   :  478.0143      0.0086      0.0084      0.0087
Summing   :  393.0148      0.0153      0.0153      0.0155
SAXPYing  :  394.7979      0.0154      0.0152      0.0155
**********************************
##################################
mcp707/double

This test was run on an FPS MCP728 running release 2.1
with 64Mbytes of matrix memory.  This machine contains
28 Intel i860 running at 40Mhz on 7 busses.  The test was
run on 7 i860s, each with its own bus.

systst2% setenv MCPCONFIG 7,1
systst2% mrun dxfer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =    1.114289999532048      hundredths of a second
 Increase the size of the arrays if this is <30 
  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment: 1040.5151      0.0046      0.0046      0.0047
Scaling   : 1036.8066      0.0046      0.0046      0.0046
Summing   :  890.4554      0.0085      0.0081      0.0090
SAXPYing  :  887.5685      0.0089      0.0081      0.0092
**********************************

From pmk@craycos.com  Fri Oct  4 15:23:45 1991
Received: from aspen.craycos.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA08824; Fri, 4 Oct 91 15:23:45 EDT
Received: from copper.craycos.com by aspen.Craycos.COM (4.1/TotalHack-4.0)
	id AA25783; Fri, 4 Oct 91 13:22:17 MDT
Received: from zymurgy.Craycos.COM by copper.craycos.com (4.0/SMI-4.0)
	id AA22080; Fri, 4 Oct 91 13:22:15 MDT
Date: Fri, 4 Oct 91 13:22:14 MDT
From: pmk@craycos.com (Peter Klausler)
Message-Id: <9110041922.AA22080@copper.craycos.com>
To: mccalpin
Subject: Re: Attainable memory bandwidth
Status: RO

> Thanks for the clarification.  i was deliberately ignoring the I/O port
> since it is not used in the floating-point kernels that I was discussing
> and since I don't understand the details well enough to talk about it!

Has to be there to sell the machine, though. I brought up the point because
your analysis was not showing the full memory bandwidth potential of CRI's
machine.

A further point to consider is that the "memory bandwidth" of a Cray-class
machine must really be specified as three distinct values:
	* Maximum processor bandwidth (CPUS * PORTS/CPU * BANDWIDTH/PORT)
	* Maximum bank bandwidth (BANKS * BANDWIDTH/BANK)
	* Maximum arbitration mechanism bandwidth (SECTIONS * BANDWIDTH/SECTION)

In practice, the minimum of these three values constrains the performance
you'll see on a real code. As a rule of thumb, total system bandwidth in
references/second should not be less than the processing rate of
floating-point results/second; this was the embarrassing imbalance of the
early CRAY-2 machines.

This view of memory system bandwidth shows why codes can get into trouble
with memory strides divisible by powers of two as small as 4 and 8. Such
strides cut down the number of sections/quadrants/octants being used, limiting
the arbitration mechanism bandwidth.

From jbs@watson.ibm.com  Fri Oct  4 19:55:27 1991
Received: from watson.ibm.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA06315; Fri, 4 Oct 91 19:55:27 EDT
Message-Id: <9110042355.AA06315@perelandra.cms.udel.edu>
Received: from YKTVMV by watson.ibm.com (IBM VM SMTP V2R1) with BSMTP id 3969;
   Fri, 04 Oct 91 19:54:06 EDT
Date: Fri, 4 Oct 91 19:54:10 EDT
From: jbs@watson.ibm.com
To: mccalpin
Subject: stream benchmark
Status: RO

         I saw your posts about your stream benchmark.  I obtained a copy
and have been playing around with it.  I ran it on a 540 (256M memory,
xlf 2.2, option -O, n=3000000).  I enclose the output.
         I saw your post in which you try to compute a theoretical rate
for the 550.  I believe you have a slightly inaccurate picture of how the
S/6000 cache works.  When a cache miss occurs on a store operation the
appropriate line is read into the cache from main memory and the store is
performed.  When a line is brought into cache it may be necessary to put
the line it replaces back in memory.  The machine keeps track of which
cache lines have been altered.  If a cache line has not been altered
it may just be overwritten.  However if the cache miss causes a line
that has been altered to be replaced the altered line must be stored
back to memory (since the stores to the altered line have not yet been
reflected to memory).  In practice buffers are used so the incoming
cache line is read in before the outgoing line is read out.  Therefore
each of your tests involves one additional read into cache.  You may
check this by altering them to overwrite one of the inputs (the alter-
ed tests should perform better).
                          James B. Shearer
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time =   227.000000000000057      hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:  145.4545       .3340       .3300       .3400
Scaling   :  129.7297       .3740       .3700       .3800
Summing   :  144.0000       .5080       .5000       .5100
SAXPYing  :  144.0000       .5090       .5000       .5100

From patrick@mozart.convex.com  Mon Oct  7 05:22:00 1991
Received: from convex.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA08367; Mon, 7 Oct 91 05:22:00 EDT
Received: from mozart.convex.com by convex.convex.com (5.61/1.35)
	id AA25025; Sun, 6 Oct 91 22:14:35 -0500
Received: by mozart.convex.com (5.64/1.28)
	id AA20897; Sun, 6 Oct 91 22:14:34 -0500
Date: Sun, 6 Oct 91 22:14:34 -0500
From: patrick@mozart.convex.com (Patrick F. McGehearty)
Message-Id: <9110070314.AA20897@mozart.convex.com>
To: mccalpin
Subject: Re:  Maximum attainable memory bandwidth
Status: RO

I only have the original version of your program, which I seem to remember
you saying had a flaw in it.  Also, I trying to juggle several deadlines
this week, so I don't know if I can get to it.  When were you thinking
of sending out the table?
- patrick

From rosenkra@c1east.convex.com  Mon Oct  7 11:03:39 1991
Received: from c1east.convex.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA09696; Mon, 7 Oct 91 11:03:39 EDT
Received: by c1east.convex.com (5.61/1.00Ic1east)
	id AA29385; Mon, 7 Oct 91 11:02:58 -0400
Date: Mon, 7 Oct 91 11:02:58 -0400
From: rosenkra@c1east.convex.com (William Rosenkranz)
Message-Id: <9110071502.AA29385@c1east.convex.com>
To: mccalpin
Subject: Re: The Great C vs. Fortran debate
Cc: rosenkra@c1east.convex.com
Status: RO

thanx for the stream code (new rev). i was just about to pick it up
from the ftp archive anyway. yes, i have followed your progress on
this with more than a fair amount of interest.

i think i mentioned about the possible hazards of using CPU time vs
wallclock and i think you responded in some fashion (but i have been
so busy lately i am not sure i read it carefully). i do remember
seing multihead CPU times (posted by my old friend andrew zachary
in darien, CT). now obviously he needed wallclock time there :-).

the other thing that naturally comes to mind is measuring non-unit
stride (either "regular" or scatter/gather) memory references with
these common kernels. it would be interesting to look at both CPU
and wallclock in those cases as well vs unit stride.

another interesting test would be to shake down the system's ability
to page with very large strides (>> cache size). but then again,
you won't find cray datapoints with real a(100000000) :-).

i look forward to more on this. and thanx again for passing this
on. i am so swamped now that i doubt that i will be able to do much
for a few weeks...

-bill
rosenkra@convex.com

From alex@Think.COM  Tue Oct  8 10:50:06 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA11810; Tue, 8 Oct 91 10:50:06 EDT
Return-Path: <alex@Think.COM>
Received: from Berlin.Think.COM by mail.think.com; Tue, 8 Oct 91 10:48:58 -0400
Received: from Zeus.Think.COM by berlin.think.com; Tue, 8 Oct 91 10:48:55 -0400
From: Alex Vasilevsky <alex@Think.COM>
Received: by zeus.think.com; Tue, 8 Oct 91 10:48:53 EDT
Date: Tue, 8 Oct 91 10:48:53 EDT
Message-Id: <9110081448.AA18730@zeus.think.com>
To: mccalpin
In-Reply-To: "John D. McCalpin"'s message of Tue, 8 Oct 91 08:51:11 EDT <9110081251.AA11383@perelandra.cms.udel.edu>
Subject:  CM2 data for stream code
Status: RO

   Date: Tue, 8 Oct 91 08:51:11 EDT
   From: "John D. McCalpin" <mccalpin@perelandra.cms.udel.edu>

   >Here is the CM-2 data for running the stream code.

   Would you mind sharing the code that you used for this?
   I am still learning about the CM-2 and would like to 
   see how you translated this....

   Thanks!
   --
   John D. McCalpin			mccalpin@perelandra.cms.udel.edu
   Assistant Professor			mccalpin@brahms.udel.edu
   College of Marine Studies, U. Del.	DELOCN::MCCALPIN (SPAN)




Here is the code.  I did the following changes to your original source: 

1. Replaced calls to function second() with calls CM timers.

2. Rewrite the realsize() routine to just return constants, there is no
need for this function to do any work really.  The CM is an IEEE machine,
so the double precision values are 8 bytes long.

3. Change the value of n

4. Wrap another loop around all expressions, otherwise the code executes
too fast to get accurate timings.

Then, I pushed this code through a vectorizer and then into CM Fortran
compiler.

First code is the non vectorized source, the second code is the output of a
vectorizer.

xxxxxxxxxxxxxxxxxx code 1 xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
* Program: Stream
* Programmer: John D. McCalpin
* Revision: 2.0, September 30,1991
*
* This program measures memory transfer rates in MB/s for simple
* computational kernels coded in Fortran.  These numbers reveal the
* quality of code generation for simple uncacheable kernels as well
* as showing the cost of floating-point operations relative to memory
* accesses.
*
* INSTRUCTIONS:
*       1) Stream requires a cpu timing function called second().
*          A sample is shown below.  This is unfortunately rather
*          system dependent.  It helps to know the granularity of the
*          timing.  The code below assumes that the granularity is
*          1/100 seconds.
*       2) Stream requires a good bit of memory to run.
*          Adjust the Parameter 'N' in the second line of the main
*          program to give a 'timing calibration' of at least 20 clicks.
*          This will provide rate estimates that should be good to
*          about 5% precision.
*       3) Compile the code with full optimization.  Many compilers
*          generate unreasonably bad code before the optimizer tightens
*          things up.  If the results are unreasonable good, on the
*          other hand, the optimizer might be too smart for me!
*       4) Mail the results to mccalpin@perelandra.cms.udel.edu
*          Be sure to include:
*               a) computer hardware model number and software revision
*               b) the compiler flags
*               c) all of the output from the test case.
*
* Thanks!
*
      PROGRAM stream
C     .. Parameters ..
      INTEGER n,ntimes
      PARAMETER (p=256,n=40000*p,ntimes=10)
C     ..
C     .. Local Scalars ..
      DOUBLE PRECISION t,t0
      INTEGER j,k,nbpw
C     ..
C     .. Local Arrays ..
      DOUBLE PRECISION a(n),b(n),c(n),maxtime(4),mintime(4),rmstime(4),
     $     times(4,ntimes)
      INTEGER bytes(4)
      CHARACTER label(4)*11
C     ..
C     .. External Functions ..
      INTEGER realsize
      EXTERNAL CM_timer_read_cm_busy,realsize
C     ..
C     .. Intrinsic Functions ..
      INTRINSIC dble,max,min,sqrt
C     ..
C     .. Data statements ..
      DATA rmstime/4*0.0D0/,mintime/4*1.0D+36/,maxtime/4*0.0D0/
      DATA label/' Assignment:',' Scaling   :',' Summing   :',
     $     ' SAXPYing  :'/
      DATA bytes/2,2,3,3/
C     ..

*       --- SETUP --- determine precision and check timing ---

      nbpw = realsize()
      
      t = 0.0D0
      call CM_timer_clear(0)
      DO 10 j = 1,n
          a(j) = 1.0D0
          b(j) = 2.0D0
          c(j) = 0.0D0
   10 CONTINUE
      call CM_timer_stop(0)
      t = CM_timer_read_cm_busy(0) - t
      PRINT *,'Timing calibration ; time = ',t*100,' hundredths',
     $  ' of a second'
      PRINT *,'Increase the size of the arrays if this is <30 ',
     $  ' and your clock precision is =<1/100 second'
      PRINT *,'---------------------------------------------------'

*       --- MAIN LOOP --- repeat test cases NTIMES times ---
      DO 60 k = 1,ntimes

         call CM_timer_clear(0)
         call CM_timer_start(0)
         do i=1,100
         DO 20 j = 1,n
            c(j) = a(j)
 20      CONTINUE
         enddo
         call CM_timer_stop(0)
         times(1,k) = CM_timer_read_cm_busy(0) 
         
         call CM_timer_clear(0)
         call CM_timer_start(0)
         do i=1,100
         DO 30 j = 1,n
            c(j) = 3.0D0*a(j)
 30      CONTINUE
         enddo
         call CM_timer_stop(0)
         times(2,k) = CM_timer_read_cm_busy(0) 

         call CM_timer_clear(0)
         call CM_timer_start(0)
         do i=1,100
         DO 40 j = 1,n
            c(j) = a(j) + b(j)
 40      CONTINUE
         enddo
         call CM_timer_stop(0)
         times(3,k) = CM_timer_read_cm_busy(0)
         call CM_timer_clear(0)
         call CM_timer_start(0)
         do i=1,100
         DO 50 j = 1,n
            c(j) = a(j) + 3.0D0*b(j)
 50      CONTINUE
         enddo
         call CM_timer_stop(0)
         times(4,k) = CM_timer_read_cm_busy(0)
 60   CONTINUE

*       --- SUMMARY ---
C*$*NOVECTORIZE
      DO 80 k = 1,ntimes
         DO 70 j = 1,4
            rmstime(j) = rmstime(j) + (times(j,k)/100.0)**2
            mintime(j) = min(mintime(j),(times(j,k)/100.0))
            maxtime(j) = max(maxtime(j),(times(j,k)/100.0))
 70      CONTINUE
 80   CONTINUE
      WRITE (*,FMT=9000)
      DO 90 j = 1,4
         rmstime(j) = sqrt(rmstime(j)/dble(ntimes))
         WRITE (*,FMT=9010) label(j),n*bytes(j)*nbpw/mintime(j)/1.0D6,
     $        rmstime(j),mintime(j),maxtime(j)
 90   CONTINUE

 9000 FORMAT (' Function',5x,'Rate (MB/s)  RMS time  Min time  Max time'
     $        )
 9010 FORMAT (a,4 (f10.4,2x))
      END

*-------------------------------------
* INTEGER FUNCTION realsize()
*
*
      INTEGER FUNCTION realsize()
      integer ndigits
      ndigits = 16
      WRITE (*,FMT='(a)') '--------------------------------------'
      WRITE (*,FMT='(1x,a,i2,a)') 'Double precision appears to have ',
     $  ndigits,' digits of accuracy'
      IF (ndigits.LE.8) THEN
          realsize = 4
      ELSE
          realsize = 8
      END IF
      WRITE (*,FMT='(1x,a,i1,a)') 'Assuming ',realsize,
     $  ' bytes per DOUBLEPRECISION word'
      WRITE (*,FMT='(a)') '--------------------------------------'
      RETURN
      END


xxxxxxxxxxxxxxxxxx code 2 xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
C     KAP/CAF             6.23 ( 7-Dec-88)    o3,r3,d3  8 Oct 1991 10:25:08
* Program: Stream
* Programmer: John D. McCalpin
* Revision: 2.0, September 30,1991
*
* This program measures memory transfer rates in MB/s for simple
* computational kernels coded in Fortran.  These numbers reveal the
* quality of code generation for simple uncacheable kernels as well
* as showing the cost of floating-point operations relative to memory
* accesses.
*
* INSTRUCTIONS:
*       1) Stream requires a cpu timing function called second().
*          A sample is shown below.  This is unfortunately rather
*          system dependent.  It helps to know the granularity of the
*          timing.  The code below assumes that the granularity is
*          1/100 seconds.
*       2) Stream requires a good bit of memory to run.
*          Adjust the Parameter 'N' in the second line of the main
*          program to give a 'timing calibration' of at least 20 clicks.
*          This will provide rate estimates that should be good to
*          about 5% precision.
*       3) Compile the code with full optimization.  Many compilers
*          generate unreasonably bad code before the optimizer tightens
*          things up.  If the results are unreasonable good, on the
*          other hand, the optimizer might be too smart for me!
*       4) Mail the results to mccalpin@perelandra.cms.udel.edu
*          Be sure to include:
*               a) computer hardware model number and software revision
*               b) the compiler flags
*               c) all of the output from the test case.
*
* Thanks!
*
      PROGRAM stream
C     .. Parameters ..
      INTEGER n,ntimes
      PARAMETER (p=256,n=40000*p,ntimes=10)
C     ..
C     .. Local Scalars ..
      DOUBLE PRECISION t,t0
      INTEGER j,k,nbpw
C     ..
C     .. Local Arrays ..
      DOUBLE PRECISION a(n),b(n),c(n),maxtime(4),mintime(4),rmstime(4),
     $     times(4,ntimes)
      INTEGER bytes(4)
      CHARACTER label(4)*11
C     ..
C     .. External Functions ..
      INTEGER realsize
      EXTERNAL CM_timer_read_cm_busy,realsize
C     ..
C     .. Intrinsic Functions ..
      INTRINSIC dble,max,min,sqrt
C     ..
C     .. Data statements ..
      DATA rmstime/4*0.0D0/,mintime/4*1.0D+36/,maxtime/4*0.0D0/
      DATA label/' Assignment:',' Scaling   :',' Summing   :',
     $     ' SAXPYing  :'/
      DATA bytes/2,2,3,3/
C     ..
 
*       --- SETUP --- determine precision and check timing ---
 
      nbpw = realsize()
 
      t = 0.0D0
      call CM_timer_clear(0)
          A = 1.0D0
          B = 2.0D0
          C = 0.0D0
      call CM_timer_stop(0)
      t = CM_timer_read_cm_busy(0) - t
      PRINT *,'Timing calibration ; time = ',t*100,' hundredths',
     $  ' of a second'
      PRINT *,'Increase the size of the arrays if this is <30 ',
     $  ' and your clock precision is =<1/100 second'
      PRINT *,'---------------------------------------------------'
 
*       --- MAIN LOOP --- repeat test cases NTIMES times ---
      DO 60 k = 1,ntimes
 
         call CM_timer_clear(0)
         call CM_timer_start(0)
         DO 20 I=1,100
            C = A
 20      CONTINUE
         call CM_timer_stop(0)
         times(1,k) = CM_timer_read_cm_busy(0)
 
         call CM_timer_clear(0)
         call CM_timer_start(0)
         DO 30 I=1,100
            C = 3.0D0 * A
 30      CONTINUE
         call CM_timer_stop(0)
         times(2,k) = CM_timer_read_cm_busy(0)
 
         call CM_timer_clear(0)
         call CM_timer_start(0)
         DO 40 I=1,100
            C = A + B
 40      CONTINUE
         call CM_timer_stop(0)
         times(3,k) = CM_timer_read_cm_busy(0)
         call CM_timer_clear(0)
         call CM_timer_start(0)
         DO 50 I=1,100
            C = A + 3.0D0 * B
 50      CONTINUE
         call CM_timer_stop(0)
         times(4,k) = CM_timer_read_cm_busy(0)
 60   CONTINUE
 
*       --- SUMMARY ---
C*$*NOVECTORIZE
      DO 80 k = 1,ntimes
         DO 70 j = 1,4
            rmstime(j) = rmstime(j) + times(j,k)**2
            mintime(j) = min(mintime(j),times(j,k))
            maxtime(j) = max(maxtime(j),times(j,k))
 70      CONTINUE
 80   CONTINUE
      WRITE (*,FMT=9000)
      DO 90 j = 1,4
         rmstime(j) = sqrt(rmstime(j)/dble(ntimes))
         WRITE (*,FMT=9010) label(j),n*bytes(j)*nbpw/mintime(j)/1.0D6,
     $        rmstime(j),mintime(j),maxtime(j)
 90   CONTINUE
 
 9000 FORMAT (' Function',5x,'Rate (MB/s)  RMS time  Min time  Max time'
     $        )
 9010 FORMAT (a,4 (f10.4,2x))
      END
C     KAP/CAF             6.23 ( 7-Dec-88)    o3,r3,d3  8 Oct 1991 10:25:08
 
*-------------------------------------
* INTEGER FUNCTION realsize()
*
*
      INTEGER FUNCTION realsize()
      integer ndigits
      ndigits = 16
      WRITE (*,FMT='(a)') '--------------------------------------'
      WRITE (*,FMT='(1x,a,i2,a)') 'Double precision appears to have ',
     $  ndigits,' digits of accuracy'
      IF (ndigits.LE.8) THEN
          realsize = 4
      ELSE
          realsize = 8
      END IF
      WRITE (*,FMT='(1x,a,i1,a)') 'Assuming ',realsize,
     $  ' bytes per DOUBLEPRECISION word'
      WRITE (*,FMT='(a)') '--------------------------------------'
      RETURN
      END

From alex@Think.COM  Tue Oct  8 11:38:56 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA11979; Tue, 8 Oct 91 11:38:56 EDT
Return-Path: <alex@Think.COM>
Received: from Berlin.Think.COM by mail.think.com; Tue, 8 Oct 91 11:37:53 -0400
Received: from Zeus.Think.COM by berlin.think.com; Tue, 8 Oct 91 11:37:51 -0400
From: Alex Vasilevsky <alex@Think.COM>
Received: by zeus.think.com; Tue, 8 Oct 91 11:37:49 EDT
Date: Tue, 8 Oct 91 11:37:49 EDT
Message-Id: <9110081537.AA18807@zeus.think.com>
To: mccalpin
In-Reply-To: "John D. McCalpin"'s message of Tue, 8 Oct 91 10:56:26 EDT <9110081456.AA11831@perelandra.cms.udel.edu>
Subject:  CM2 data for stream code
Status: RO

   Date: Tue, 8 Oct 91 10:56:26 EDT
   From: "John D. McCalpin" <mccalpin@perelandra.cms.udel.edu>

   Thanks for the code.  I had heard that you folks were using a KAP
   vectorizer in-house, and it is not surprising that it would work
   on something this simple!
   --
   John D. McCalpin			mccalpin@perelandra.cms.udel.edu
   Assistant Professor			mccalpin@brahms.udel.edu
   College of Marine Studies, U. Del.	DELOCN::MCCALPIN (SPAN)


No problem.  We use it internally, but very little.  

	-Alex V.

From mash@mips.com  Fri Oct 11 14:44:56 1991
Received: from spim.mips.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA00868; Fri, 11 Oct 91 14:44:56 EDT
Received: from winchester.mips.com by spim.mips.com via SMTP (5.61.15/2.9)
	id AA29487; Fri, 11 Oct 91 11:44:01 -0700
Received: by winchester.mips.com (5.61/Relay-2.9) 
	id AA05707; Fri, 11 Oct 91 11:44:05 -0700
Date: Fri, 11 Oct 91 11:44:05 -0700
From: mash@mips.com (John Mashey)
Message-Id: <9110111844.AA05707@winchester.mips.com>
To: mccalpin
Subject: Re: Floating Point Systems
Status: RO

thanx.  Are any of the numbers actually running on the SPARC part,
or on the other pieces?

From lupienj@hpwarq.wal.hp.com  Wed Oct 16 16:37:24 1991
Received: from relay.hp.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA01038; Wed, 16 Oct 91 16:37:24 EDT
Received: from hpwala.wal.hp.com by relay.hp.com with SMTP
	(16.6/15.5+IOS 3.13) id AA19950; Wed, 16 Oct 91 13:36:46 -0700
Received: from hpwadac.wal.hp.com by hpwala.wal.hp.com with SMTP
	(15.11/15.5+IOS 3.22) id AA19593; Wed, 16 Oct 91 16:38:22 edt
Received: by hpwarq.HP.COM; Wed, 16 Oct 91 16:36:12 edt
Date: Wed, 16 Oct 91 16:36:12 edt
From: John Lupien <lupienj@hpwarq.wal.hp.com>
Message-Id: <9110162036.AA11674@hpwarq.HP.COM>
To: mccalpin
Subject: Re: Attainable memory bandwidth
Newsgroups: comp.arch,comp.benchmarks
In-Reply-To: <MCCALPIN.91Sep30143808@pereland.cms.udel.edu>
References: <MCCALPIN.91Sep27212326@pereland.cms.udel.edu>
Organization: Hewlett Packard, Waltham, Mass
Cc: 
Status: RO

In article <MCCALPIN.91Sep30143808@pereland.cms.udel.edu> you write:
>By the way, I have not heard from anyone with an HP-720 or HP-730.
>Don't all you HP fans out there want to show off?

While we might want to show off, it's not at all clear that we are allowed to.
If you would like to send me a program to run, I can ask if we can tell you
the results.... (sounding like a gutless marine bottom feeder...)

-- 
---
John R. Lupien
lupienj@hpwarq.hp.com

From lupienj@hpwarq.wal.hp.com  Wed Oct 16 17:22:59 1991
Received: from relay.hp.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA01149; Wed, 16 Oct 91 17:22:59 EDT
Received: from hpwala.wal.hp.com by relay.hp.com with SMTP
	(16.6/15.5+IOS 3.13) id AA21263; Wed, 16 Oct 91 14:22:24 -0700
Received: from hpwadac.wal.hp.com by hpwala.wal.hp.com with SMTP
	(15.11/15.5+IOS 3.22) id AA21614; Wed, 16 Oct 91 17:24:01 edt
Received: by hpwarq.HP.COM; Wed, 16 Oct 91 17:21:51 edt
Message-Id: <9110162121.AA11835@hpwarq.HP.COM>
From: lupienj@hpwarq.wal.hp.com (John Lupien)
Date: Wed, 16 Oct 91 17:21:32 EDT
X-Server: hpwadac
X-Citing: Celtics!
X-Cargo: steamed snails
X-Mailer: Mail User's Shell (6.5.6 6/30/89)
To: "John D. McCalpin" <mccalpin>
Subject: Re: Attainable memory bandwidth
Status: RO

> 33 MHz cycles (8 for the transfer and n for latency), then
> the bandwidth is 
> 	53.3 MB/s = 64 bytes/(n+8 cycles) * 33 Million cycles/sec 
> which implies a latency of about n=32 cycles.
> For the 720
> 	43.6 MB/s = 64 bytes/(n+8 cycles) * 25 Million cycles/sec 
> which implies a latency of about n=29 cycles.
> Both of these numbers seem unreasonably high.  
> Am I completely misunderstanding the way the memory system on the
> HP works?

Sorry, I just write software.... I'm lucky if I get to see "how fast it runs"
on a 700 machine (it's amazingly fast for our applications).
I'll forward your comments to someone who may havea better answer.


-- 
---
John R. Lupien
lupienj@hpwarq.hp.com

From broadley@schneider3.lrdc.pitt.edu  Thu Oct 17 17:39:51 1991
Received: from schneider3.lrdc.pitt.edu by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA04253; Thu, 17 Oct 91 17:39:51 EDT
Received: by schneider3.lrdc.pitt.edu (5.57/Ultrix3.0-C)
	id AA04294; Thu, 17 Oct 91 17:39:13 -0400
Date: Thu, 17 Oct 91 17:39:13 -0400
From: broadley@schneider3.lrdc.pitt.edu (Bill Broadley)
Message-Id: <9110172139.AA04294@schneider3.lrdc.pitt.edu>
To: mccalpin
Subject: Mem transfer benchmark
Status: RO

	Did you get any numbers on the SGI 4d machines?  Like the 480??
			-Bill
			Broadley@schneider3.lrdc.pitt.edu

From patrick@mozart.convex.com  Sat Oct 19 17:41:30 1991
Received: from convex.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA00316; Sat, 19 Oct 91 17:41:30 EDT
Received: from mozart.convex.com by convex.convex.com (5.61/1.35)
	id AA06366; Fri, 18 Oct 91 16:14:26 -0500
Received: by mozart.convex.com (5.64/1.28)
	id AA06659; Fri, 18 Oct 91 16:14:25 -0500
Date: Fri, 18 Oct 91 16:14:25 -0500
From: patrick@mozart.convex.com (Patrick F. McGehearty)
Message-Id: <9110182114.AA06659@mozart.convex.com>
To: mccalpin
Subject: Re:  Maximum attainable memory bandwidth
Status: RO

Sorry for my lack of response.  We are getting ready to ship an
compiler release, new products, etc, etc :-)  But enough excuses.

The real reason for the delay is that I decided to run it by marketing for
their pro forma approval before sending results for a publishable
benchmark, as a courtesy.  Then I found out that the appropriate people were
out of town (customer visits, SuperComputing '91, etc) and won't be back for
a little while.  Having started this route, I feel I have to wait for a
acknowledgement, so I have no offical results to report.  I will let you
know when I hear something, but I don't expect anything for another week
at least.

Unofficially, compiling fc -O2 generates real*4 vector code,
while fc -O2 -pd8 generates real*8 vector code.  [The -pd8 says
set default real/integer length to 8].  On the C2 and C3200
product line, there is almost a factor of two difference between
the measured results for these two options, since the clock rate
is the same for both, but real*4 works on 4 bytes at a time while
real*8 works on 8 bytes at a time.  Preliminary results show no
major surprises, but there are still plenty of holes in the 
chart of C3200/C3400/C3800/ 1-4 processors.

- patrick

From patrick@mozart.convex.com  Mon Oct 28 14:57:20 1991
Received: from convex.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA11438; Mon, 28 Oct 91 14:57:20 EST
Received: from mozart.convex.com by convex.convex.com (5.61/1.35)
	id AA18834; Mon, 28 Oct 91 13:59:15 -0600
Received: by mozart.convex.com (5.64/1.28)
	id AA11185; Mon, 28 Oct 91 13:59:13 -0600
Date: Mon, 28 Oct 91 13:59:13 -0600
From: patrick@mozart.convex.com (Patrick F. McGehearty)
Message-Id: <9110281959.AA11185@mozart.convex.com>
To: mccalpin
Subject: Re:  Maximum attainable memory bandwidth
Status: RO

I've gotten an okay from marketing, and just want to collect a couple
of more data points for some different configurations tonight.  I'll
try to send you data tomorrow.  Thanks for your patience.
- patrick

From patrick@mozart.convex.com  Tue Oct 29 14:50:54 1991
Received: from convex.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA12994; Tue, 29 Oct 91 14:50:54 EST
Received: from mozart.convex.com by convex.convex.com (5.61/1.35)
	id AA22750; Tue, 29 Oct 91 13:51:04 -0600
Received: by mozart.convex.com (5.64/1.28)
	id AA09236; Tue, 29 Oct 91 13:51:02 -0600
Date: Tue, 29 Oct 91 13:51:02 -0600
From: patrick@mozart.convex.com (Patrick F. McGehearty)
Message-Id: <9110291951.AA09236@mozart.convex.com>
To: mccalpin
Subject: Long awaited Max memory bandwidth results
Status: RO

Here are the outputs from a series of runs.  They include
1 and 2 processor tests for real*4 and real*8 data on
the C3200 series and C3400 series (total of 8 data results).
I could not get access to a fully configured C3800 to run this benchmark,
as the few we have are dedicated to paying customer benchmarks and final
product development.  Next quarter I expect to have some impressive
results (clock speed diffs suggest a 2.4 times increase in speed).

As we have discussed before, these benchmarks don't properly represent
our compiler's ability to reuse vector data for many common loops such
as linpack and matrix multiply to obtain a MUL-ADD for each memory
reference.  However, they do measure one important aspect of total
system performance.

I separated each set of results with a line of +++'s and a one line
description of the machine configuration which it ran on.

- Patrick McGehearty (patrick@convex.com)

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3210 with 32 way interleave, real = 4 bytes
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    4.399200     hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0435      0.0428      0.0442
Scaling   :    0.0437      0.0434      0.0440
Summing   :    0.0554      0.0551      0.0557
SAXPYing  :    0.0554      0.0551      0.0558


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:   91.9481     91.9782     93.4448     90.4466
Scaling   :   91.5923     91.6300     92.1766     90.8945
Summing   :  108.3314    108.3740    108.8673    107.8128
SAXPYing  :  108.2118    108.2583    108.9701    107.4498

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3210 with 32 way interleave, real = 8 bytes
Test #1 Failed = picalc=piexact
Apparently Single=Double Precision
Proceeding to Test #2
 
--------------------------------------
 Single precision appears to have 16 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    4.45790000000000      hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0447      0.0444      0.0451
Scaling   :    0.0449      0.0446      0.0453
Summing   :    0.0666      0.0662      0.0671
SAXPYing  :    0.0669      0.0664      0.0677


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:  178.8493    178.9316    180.0950    177.2971
Scaling   :  178.0718    178.1687    179.4245    176.5576
Summing   :  180.0483    180.1261    181.3565    178.7603
SAXPYing  :  179.3529    179.4383    180.8209    177.2473

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3220 with 32 way interleave, real = 4 bytes
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    2.309900     hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0226      0.0224      0.0230
Scaling   :    0.0227      0.0225      0.0236
Summing   :    0.0292      0.0284      0.0324
SAXPYing  :    0.0291      0.0287      0.0294


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:  177.0076    177.1239    178.6831    173.7544
Scaling   :  175.8365    175.9640    177.5331    169.1691
Summing   :  205.5435    205.5453    210.9705    184.9742
SAXPYing  :  206.1679    206.3336    208.7541    203.7487

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3220 with 32 way interleave, real = 8 bytes
Test #1 Failed = picalc=piexact
Apparently Single=Double Precision
Proceeding to Test #2
 
--------------------------------------
 Single precision appears to have 16 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    2.39510000000000      hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0243      0.0239      0.0252
Scaling   :    0.0242      0.0240      0.0248
Summing   :    0.0363      0.0356      0.0396
SAXPYing  :    0.0360      0.0357      0.0369


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:  329.3360    329.6101    334.2665    318.0662
Scaling   :  329.6278    329.9916    333.0281    322.5416
Summing   :  330.3847    330.5466    337.2776    302.8085
SAXPYing  :  332.7483    333.0980    336.4832    325.0007

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3410 with 32 way interleave, real = 4 bytes
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    2.812000     hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0262      0.0258      0.0270
Scaling   :    0.0269      0.0265      0.0276
Summing   :    0.0406      0.0382      0.0491
SAXPYing  :    0.0395      0.0389      0.0414


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:  152.7073    152.8022    155.1109    148.3459
Scaling   :  148.4775    148.5742    150.9206    144.7491
Summing   :  148.1353    147.6949    157.0846    122.0827
SAXPYing  :  151.6438    151.7295    154.4086    144.8397

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3410 with 32 way interleave, real = 8 bytes
Test #1 Failed = picalc=piexact
Apparently Single=Double Precision
Proceeding to Test #2
 
--------------------------------------
 Single precision appears to have 16 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    5.60450000000000      hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0466      0.0450      0.0482
Scaling   :    0.0469      0.0464      0.0475
Summing   :    0.0687      0.0670      0.0761
SAXPYing  :    0.0689      0.0672      0.0765


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:  171.6270    171.7024    177.5923    165.9476
Scaling   :  170.5793    170.6833    172.5216    168.4388
Summing   :  174.5774    174.5629    179.0243    157.6086
SAXPYing  :  174.2757    174.2588    178.4493    156.7972

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3420 with 32 way interleave, real = 4 bytes
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    1.751800     hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0170      0.0165      0.0179
Scaling   :    0.0172      0.0167      0.0180
Summing   :    0.0232      0.0228      0.0254
SAXPYing  :    0.0237      0.0231      0.0270


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:  235.0922    235.2967    242.6743    224.0268
Scaling   :  232.0791    232.3044    238.8061    222.5933
Summing   :  258.4247    258.6258    263.7243    236.6210
SAXPYing  :  252.9180    252.9793    259.9990    221.8773

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Test run on C3420 with 32 way interleave, real = 8 bytes
Test #1 Failed = picalc=piexact
Apparently Single=Double Precision
Proceeding to Test #2
 
--------------------------------------
 Single precision appears to have 16 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
Timing calibration ; time =    2.90190000000000      hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
Or ignore variations in timing and just look at
 totals
---------------------------------------------------
Function      RMS time    Min time    Max time
Assignment:    0.0298      0.0277      0.0338
Scaling   :    0.0298      0.0293      0.0316
Summing   :    0.0414      0.0396      0.0492
SAXPYing  :    0.0406      0.0400      0.0419


Memory transfer rates in MB/s

Function       Total       RMS         Best        Worst
Assignment:  268.6195    268.5465    289.1531    236.6094
Scaling   :  267.9403    268.1232    273.1121    253.1405
Summing   :  290.0807    289.7490    302.7016    243.8826
SAXPYing  :  295.0962    295.3977    300.3003    286.6767

From tighe@convex1.convex.com  Fri Nov  1 10:42:21 1991
Received: from convex.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17604; Fri, 1 Nov 91 10:42:21 EST
Received: from convex1.convex.com by convex.convex.com (5.61/1.35)
	id AA06214; Fri, 1 Nov 91 09:42:46 -0600
Received: by convex1.convex.com (5.64/1.28)
	id AA12593; Fri, 1 Nov 91 09:42:45 -0600
From: tighe@convex1.convex.com (Mike Tighe)
Message-Id: <9111011542.AA12593@convex1.convex.com>
Subject: Question
To: mccalpin
Date: Fri, 1 Nov 91 9:42:44 CST
Organization: Convex Computer Corporation, Richardson, Texas
X-Mailer: ELM [version 2.3 PL11]
Status: RO

You write:

>                     Bytes          Bandwidth (MB/s)                     
>  achine             /word      Copy     SSCAL       Sum     SAXPY       
> ---------------     -----  --------  --------  --------  --------       
> Convex C3420 2 cpu      8     289.1     273.1     302.7     300.3       
> Convex C3410 1 cpu      8     177.6     172.5     179.0     178.4       

How did you obtain 3400 rates? I ask because that is a new machine, and I
didn't think that there were many out there yet in customer hands.

-Mike

From don@mars.dgrc.doc.ca  Fri Nov  1 13:37:10 1991
Received: from dgbt.doc.ca by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17835; Fri, 1 Nov 91 13:37:10 EST
Received: by dgbt.doc.ca (5.57/smail2.5/12-02-88)
	id AA04182; Fri, 1 Nov 91 13:37:17 EST
Received: from jack.dgrc.doc.ca by mars.dgrc.doc.ca.dgrc.doc.ca (4.1/SMI-4.1)
	id AA03497; Fri, 1 Nov 91 13:31:18 EST
Date: Fri, 1 Nov 91 13:31:18 EST
From: don@mars.dgrc.doc.ca (Donald McLachlan)
Message-Id: <9111011831.AA03497@mars.dgrc.doc.ca.dgrc.doc.ca>
To: mccalpin
Subject: RE: memory bandwidth
Status: RO

I suspect the results you have listed for a Sun SS1 may be for a Sun SS1+.

I just compiled stream.f with no optimizations, and got the following MBytes/s
6.2500
5.0314
6.5217
5.7143

I then compiled with -Bstatic -cg89 -dalign -f -fast -O4 -native and got
10.3897
 9.1954
10.1695
10.1695

I then went into single user mode and got
10.5263
 9.3023
10.2564
10.2564

Just want to keep things honest, Don

From tighe@convex1.convex.com  Fri Nov  1 14:40:14 1991
Received: from convex.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA17886; Fri, 1 Nov 91 14:40:14 EST
Received: from convex1.convex.com by convex.convex.com (5.61/1.35)
	id AA23467; Fri, 1 Nov 91 13:40:34 -0600
Received: by convex1.convex.com (5.64/1.28)
	id AA21683; Fri, 1 Nov 91 13:40:33 -0600
From: tighe@convex1.convex.com (Mike Tighe)
Message-Id: <9111011940.AA21683@convex1.convex.com>
Subject: Re:  Question
To: mccalpin (John D. McCalpin)
Date: Fri, 1 Nov 91 13:40:33 CST
In-Reply-To: <9111011706.AA17716@perelandra.cms.udel.edu>; from "John D. McCalpin" at Nov 1, 91 12:06 pm
Organization: Convex Computer Corporation, Richardson, Texas
X-Mailer: ELM [version 2.3 PL11]
Status: RO

John D. McCalpin writes:

>You should point out the unexpectedly bad results of the HP9000/720 and 
>HP9000/730 to your marketing department!  The results are only about
>1/4 of what HP implies in their literature.

Funny you should mention that (are you psychic?). I have an xterm open to a
720 machine and preparing to run some benchmarks through it.

Thanks for the info. I have obtained the source for the test from Patrick,
and will try to duplicate the results on some other machines.

-Mike




From mfriedel@monolith.rmNUG.ORG  Sat Nov  2 23:55:46 1991
Received: from nugget.rmNUG.ORG by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA18820; Sat, 2 Nov 91 23:55:46 EST
Received: by nugget with UUCP id AA12929
  (5.65c/IDA-1.4.4 for mccalpin@perelandra.cms.udel.edu); Sat, 2 Nov 1991 21:51:31 -0700
Received: by  monolith.DaVinci.COM  (NeXT-1.0 (From Sendmail 5.52)/DaVinci-sub-$Revision: 1.2 $)
	id AA09255; Sat, 2 Nov 91 21:52:13 MST
Date: Sat, 2 Nov 91 21:52:13 MST
From: mfriedel@monolith.rmnug.org (Michael Friedel)
Message-Id: <9111030452.AA09255@ monolith.DaVinci.COM >
Received: by NeXT Mailer (1.63)
To: mccalpin
Subject: Memory Bandwidth
Status: RO

I converted your stream tests to C and ran them on my NeXT.

Here ar ethe results

Hardware:
NeXTStation Color, 12Meg
NeXTStep 2.1, with BSD 4.3 kernel

Compiler flags
-O 


I tried all kinds of combinations, but they didn'y make a  
difference. Note: Although the station has a 68040 it uses  
68030 code. 


Results

For double precision

Timing calibration: Time = 132.7215 hundredths of a second
Increase the size of the arrays if this is < 30 and your 

clock precision is <= 1/100 second
Number of bytes per word = 8

Function         Rate (MB/s)    RMS time   Min time  Max time
Assignment :      14.2369        0.1079     0.3372    0.3582
Scaling    :      11.4952        0.1357     0.4176    0.4599
Summing    :      16.4071        0.1430     0.4388    0.4949
SAXPYing   :      16.1239        0.1437     0.4465    0.4846

And for single precision

Timing calibration: Time = 60.2503 hundredths of a second
Increase the size of the arrays if this is < 30 and your 

clock precision is <= 1/100 second
Number of bytes per word = 4

Function         Rate (MB/s)    RMS time   Min time  Max time
Assignment:       13.4077        0.0568     0.1790    0.1802
Scaling   :        7.6574        0.0995     0.3134    0.3157
Summing   :       10.5189        0.1085     0.3422    0.3438
SAXPYing  :        8.4488        0.1350     0.4261    0.4275


Mike

From uunet.UU.NET!meaddata!meaddata.com!chuckg  Tue Nov  5 13:15:38 1991
Received: from relay2.UU.NET by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA22647; Tue, 5 Nov 91 13:15:38 EST
Received: from uunet.uu.net (via LOCALHOST.UU.NET) by relay2.UU.NET with SMTP 
	(5.61/UUNET-internet-primary) id AA28860; Tue, 5 Nov 91 13:16:25 -0500
Received: from meaddata.UUCP by uunet.uu.net with UUCP/RMAIL
	(queueing-rmail) id 131509.11401; Tue, 5 Nov 1991 13:15:09 EST
Received: from herman.meaddata.com by meaddata.com (4.1/SMI-4.1)
	id AA02187; Tue, 5 Nov 91 13:12:45 EST
Received: by herman.meaddata.com (4.1/SMI-4.1)
	id AA07817; Tue, 5 Nov 91 13:12:44 EST
Date: Tue, 5 Nov 91 13:12:44 EST
From: chuckg@meaddata.com (Chuck Greenwald)
Message-Id: <9111051812.AA07817@herman.meaddata.com>
To: mccalpin (John D. McCalpin)
Subject: Re: Memory Bandwidth Table
Status: RO

I'ld like to run your benchmark on our IBM mainframes, but don't have FTP access.
Could you mail me your Fortran code?  Thanks!

-------------------------------------------------------------------------
Chuck Greenwald                     |   chuckg@meaddata.com  
Mead Data Central                   |   (513) 865-1020

From alz@grumpy.cray.com  Tue Nov  5 13:16:07 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA22653; Tue, 5 Nov 91 13:16:07 EST
Received: from dopey.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6k)
	id AA13764; Tue, 5 Nov 91 12:16:44 CST
Received: by dopey.cray.com
	id AA00498; 4.1/CRI-4.4; Tue, 5 Nov 91 13:16:40 EST
Date: Tue, 5 Nov 91 13:16:40 EST
From: alz@grumpy.cray.com (Andrew Zachary)
Message-Id: <9111051816.AA00498@dopey.cray.com>
To: mccalpin
Subject: Re: Memory Bandwidth Table
Status: RO

John,

In looking throught your latest posting of the memory bandwidth table,
I was struck by the numbers for the CM2.  In particular, the numbers
you report for the SAXPY and SUM show speeds greater than 80 Gbytes/sec.
The numbers I have indicate that the limiting memory bandwidth on
a fully configure CM2 will be

	4 bytes/sec/processor * 2000 processors * 10 Mhz = 80 Gbytes/sec

This number should be a "guaranteed not to exceed" performance indicator.
How, then, did TMC produce a larger number?  Did they use arrays small
enough to fit inside the registers on the Wyteks?  Did they "tweek" the
machine to improve the clock speed?  Is there a cache on the Wyteks? 
Do you have any other thoughts about how TMC exceeded their speed-of-light
limitation?

Thanks,
Andrew Zachary
alz@grumpy.cray.com

From alex@Think.COM  Tue Nov  5 14:42:49 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA23005; Tue, 5 Nov 91 14:42:49 EST
Return-Path: <alex@Think.COM>
Received: from Berlin.Think.COM by mail.think.com; Tue, 5 Nov 91 14:43:14 -0500
Received: from Zeus.Think.COM by berlin.think.com; Tue, 5 Nov 91 14:43:12 -0500
From: Alex Vasilevsky <alex@Think.COM>
Received: by zeus.think.com; Tue, 5 Nov 91 14:43:10 EST
Date: Tue, 5 Nov 91 14:43:10 EST
Message-Id: <9111051943.AA10504@zeus.think.com>
To: mccalpin
In-Reply-To: "John D. McCalpin"'s message of Tue, 5 Nov 91 13:51:44 EST <9111051851.AA22784@perelandra.cms.udel.edu>
Subject: memory bandwidth
Status: RO

   Date: Tue, 5 Nov 91 13:51:44 EST
   From: "John D. McCalpin" <mccalpin@perelandra.cms.udel.edu>

   I was looking over the memory bandwidth results that you provided
   for the CM-2, and had a question I hope you could answer.
   The Copy and Scale operations ran at about 56 GB/s on the 64-K machine.
   This implies a memory bandwidth limitation of
	   4 Bytes/clock/fpu * 2048 fpus * 7 MHz = 57.3 GB/s
   The Sum and Triad operations ran at about 80 GB/s, which implies 
   that there is more data bandwidth available.  Do the fpus have more
   than one independent 32-bit data path?
   Thanks for any explanation!
   --
   John D. McCalpin			mccalpin@perelandra.cms.udel.edu
   Assistant Professor			mccalpin@brahms.udel.edu
   College of Marine Studies, U. Del.	DELOCN::MCCALPIN (SPAN)


No.  It is just that the Fortran compiler is very clever and generates the
best possible sequence of code for these cases.  By looking at the timings,
even though the machine is doing an extra load, the timing is not much
different from the code doing 2 loads, the difference is about a 5%.  I see
a very similar thing going on the Cray.

CM2 64K
--------
gorka(test)% stream
-------------------------------------
Double precision appears to have 16 digits of accuracy
Assuming 8 bytes per DOUBLEPRECISION word
-------------------------------------
Calibrating CM timer...Done. CM speed = 7.00 MHz
Timing calibration ; time =   143.324911594391 hundredths of a second
Increase the size of the arrays if this is <30  and your clock precision is =<1
/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment  56555.9654       .0231       .0231       .0231
Scaling     56555.9654       .0231       .0231       .0231
Summing     81244.0132       .0242       .0242       .0242
SAXPYing    80593.2540       .0244       .0244       .0244



Here are the two DO loops and the vector code that our compiler generates:

First DO loop:

         DO 40 j = 1,n
            c(j) = a(j) + b(j)
 40      CONTINUE


Code generated by compiler, every instruction is a vector instruction here.

procedure _stream_pe_code_3
L1$_stream_pe_code_3:
                                                                    
	popq       aC2                                              
	popa       SP                                               
	# Get address of A
	popa       aP2                                              
	# Get address of B
	popa       aP3                                              
	# Get address of C
	popa       aP4

L2$_stream_pe_code_3:

	dflodv     [aP3+0]2++ aV0
	# "stream_d.fcm" line 102
	# C = A + B
	dfaddv     [aP2+0]2++ aV0 aV1
	dfstrv     aV1 [aP4+0]2++    
	jnz        aC2 L2$_stream_pe_code_3
end

Second DO loop:

         DO 50 j = 1,n
            c(j) = a(j) + 3.0D0*b(j)
 50      CONTINUE


Code generated by compiler, every instruction is a vector instruction here.

procedure _stream_pe_code_4
L1$_stream_pe_code_4:

	popq       aC2
	popa       SP 
	# Get address of A
	popa       aP2
	# Get address of B
	popa       aP3  
	# Get address of C
	popa       aP4  
	# "stream_d.fcm" line 110
	# C = A + 3.0D0*B
	dflodc     $3.000000000000000000d+00       aS28
                                                       

L2$_stream_pe_code_4:

	dflodv     [aP2+0]2++ aV0
	# "stream_d.fcm" line 110
	# C = A + 3.0D0*B
	dfmuladdv  aS28 [aP3+0]2++ aV1 aV0 aV1
	dfstrv     aV1 [aP4+0]2++             
	jnz        aC2 L2$_stream_pe_code_4   

end

From alz@grumpy.cray.com  Tue Nov  5 15:19:44 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA23093; Tue, 5 Nov 91 15:19:44 EST
Received: from dopey.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6k)
	id AA19775; Tue, 5 Nov 91 14:20:04 CST
Received: by dopey.cray.com
	id AA00531; 4.1/CRI-4.4; Tue, 5 Nov 91 15:19:59 EST
Date: Tue, 5 Nov 91 15:19:59 EST
From: alz@grumpy.cray.com (Andrew Zachary)
Message-Id: <9111052019.AA00531@dopey.cray.com>
To: mccalpin
Subject: Re: Memory Bandwidth Table
In-Reply-To: Mail from 'mccalpin@perelandra.cms.udel.edu'
      dated: Tue, 5 Nov 91 15:03:32 EST
Status: RO

John,

Would it be possible for you to follow up with TMC and find out how
they performed their magic?  From your numbers, the maximum memory
bandwidth for a CM2 should be

	4 Bytes/sec/process * 2048 processors * 7 Mhz = 57.3 Gbytes/sec

even for SAXPY's or SUM's.

Thanks,
Andrew Zachary


From sandee@Think.COM  Tue Nov  5 18:24:16 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA23360; Tue, 5 Nov 91 18:24:16 EST
Return-Path: <sandee@Think.COM>
Received: from Lugh.Think.COM by mail.think.com; Tue, 5 Nov 91 18:24:58 -0500
From: Daan Sandee <sandee@Think.COM>
Received: by lugh.think.com (4.1/Think-1.0C)
	id AA29761; Tue, 5 Nov 91 18:24:57 EST
Date: Tue, 5 Nov 91 18:24:57 EST
Message-Id: <9111052324.AA29761@lugh.think.com>
To: mccalpin
Subject: Re: vector code
Cc: sandee@Think.COM
Status: RO

>L2$_stream_pe_code_3:
>	dflodv     [aP3+0]2++ aV0
>	# C = A + B
>	dfaddv     [aP2+0]2++ aV0 aV1
>	dfstrv     aV1 [aP4+0]2++    
>	jnz        aC2 L2$_stream_pe_code_3	! this is an iteration counter

Okay, I've calmed down, and now I *do* recognize the CM-2.
Every instruction is a macro which expands into actual CMIS code.
So, [above]
 load into register file using memory address at P3 strided by 2 (slices) and
   using register file address V0;
 load (ditto, P2) and add to register contents using RF address V0 > V1 ;
 store (ditto, P4) from RF at RF address V1.

>	# C = A + 3.0D0*B
>	dflodc     $3.000000000000000000d+00       aS28
>L2$_stream_pe_code_4:
>	dflodv     [aP2+0]2++ aV0
>	# C = A + 3.0D0*B
>	dfmuladdv  aS28 [aP3+0]2++ aV1 aV0 aV1
>	dfstrv     aV1 [aP4+0]2++             
>	jnz        aC2 L2$_stream_pe_code_4   

This appears to be the same thing except the Weitek chip does a multiply-
with-constant (in one hardware instruction ; I don't know the op codes.)

So what has this to do with memory bandwidth ? Nothing. In both cases
it loads 3 data words per 6 cycles ; in the second case it does two flops
in stead of one, per 6 cycles (and per output word). I'd have to go to the
chip manual to see if there is anything to be got out of the timing specs.


From don@mars.dgrc.doc.ca  Wed Nov  6 10:10:43 1991
Received: from [192.12.98.7] by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA23962; Wed, 6 Nov 91 10:10:43 EST
Received: by dgbt.doc.ca (5.57/smail2.5/12-02-88)
	id AA16158; Wed, 6 Nov 91 10:11:20 EST
Received: from jack.dgrc.doc.ca by mars.dgrc.doc.ca.dgrc.doc.ca (4.1/SMI-4.1)
	id AA01988; Wed, 6 Nov 91 10:04:57 EST
Date: Wed, 6 Nov 91 10:04:57 EST
From: don@mars.dgrc.doc.ca (Donald McLachlan)
Message-Id: <9111061504.AA01988@mars.dgrc.doc.ca.dgrc.doc.ca>
To: mccalpin
Subject: RE: memory Bandwidth
Status: RO


Following up on our earlier correspondance ...

I have redone my testing of stream_s.f and stream_d.f and still believe
the numbers shown in your table are incorrect for a Sun SS1.

I have used script which explains all the ''s.

Don

... Here are my results ...

Script started on Wed Nov  6 09:57:24 1991
jack don> f77 -cg89 -dalign -f -fast -O4 -native stream_s.f -o stream_s
f77: Warning: -O4 overwrites previously set optimization level of -O2
stream_s.f:
 MAIN stream:
        second:
        realsize:
        dummy:
jack don> stream_s
--------------------------------------
 Single precision appears to have  7 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
Timing calibration ; time =     177.000 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   10.5263      0.3850      0.3800      0.3900
Scaling   :    9.3023      0.4370      0.4300      0.4400
Summing   :   10.1695      0.5960      0.5900      0.6000
SAXPYing  :   10.1695      0.5991      0.5900      0.6200
 Note: this program was linked with -fast or -fnonstd 
 and so may have produced nonstandard floating-point results. 
 Sun's implementation of IEEE arithmetic is discussed in 
 the Numerical Computation Guide.
jack don> f77 -cg89 -dalign -f -fast -O4 -native stream_d.f -o stream_d
f77: Warning: -O4 overwrites previously set optimization level of -O2
stream_d.f:
 MAIN stream:
        second:
        realsize:
        dummy:
jack don> stream_d
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
Timing calibration ; time =     210.99999696016 hundredths of a second
Increase the size of the arrays if this is <30 
 and your clock precision is =<1/100 second
---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment:   12.6316      0.3900      0.3800      0.4000
Scaling   :   12.0000      0.4030      0.4000      0.4100
Summing   :   12.6316      0.5740      0.5700      0.5800
SAXPYing  :   11.4286      0.6350      0.6300      0.6400
 Note: this program was linked with -fast or -fnonstd 
 and so may have produced nonstandard floating-point results. 
 Sun's implementation of IEEE arithmetic is discussed in 
 the Numerical Computation Guide.
jack don> ^D
script done on Wed Nov  6 10:07:24 1991

From don@mars.dgrc.doc.ca  Wed Nov  6 12:31:36 1991
Received: from dgbt.doc.ca by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA24271; Wed, 6 Nov 91 12:31:36 EST
Received: by dgbt.doc.ca (5.57/smail2.5/12-02-88)
	id AA16628; Wed, 6 Nov 91 12:32:16 EST
Received: by mars.dgrc.doc.ca.dgrc.doc.ca (4.1/SMI-4.1)
	id AA02199; Wed, 6 Nov 91 12:25:53 EST
Date: Wed, 6 Nov 91 12:25:53 EST
From: don@mars.dgrc.doc.ca (Donald McLachlan)
Message-Id: <9111061725.AA02199@mars.dgrc.doc.ca.dgrc.doc.ca>
To: mccalpin
Subject: RE: memory Bandwidth
Status: RO

I am curious, which version of SunOS and f77 do you have?
I am using SunOS 4.1.1a, and F77-1.4.

Which optimizer options did you use?

I feel as though there must be something to explain the difference then.

12.9032 vs 10.3897 MB/s = 1.24192228842026237523 looks surprising close to
20      vs 16      MHz  = 1.25

And it is only the clock speed that differenciates the SS1 and SS1+.

Don

From don@mars.dgrc.doc.ca  Wed Nov  6 13:07:26 1991
Received: from dgbt.doc.ca by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA24389; Wed, 6 Nov 91 13:07:26 EST
Received: by dgbt.doc.ca (5.57/smail2.5/12-02-88)
	id AA16775; Wed, 6 Nov 91 13:08:09 EST
Received: by mars.dgrc.doc.ca.dgrc.doc.ca (4.1/SMI-4.1)
	id AA02380; Wed, 6 Nov 91 13:01:45 EST
Date: Wed, 6 Nov 91 13:01:45 EST
From: don@mars.dgrc.doc.ca (Donald McLachlan)
Message-Id: <9111061801.AA02380@mars.dgrc.doc.ca.dgrc.doc.ca>
To: mccalpin
Subject: RE: memory Bandwidth
Status: RO

If you haven't upgraded, then your results are even more interesting since
Sun claims f77-1.4 is about 10% faster than f77-1.3.

Well it looks as though we will not be able to resolve this. Anyway, I
found the results of your benchmark interesting.

Don

From dik@cwi.nl  Thu Nov  7 17:35:41 1991
Received: from charon.cwi.nl by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA26757; Thu, 7 Nov 91 17:35:41 EST
Received: by charon.cwi.nl with SMTP; Thu, 7 Nov 1991 23:36:32 +0100
Received: by paring.cwi.nl ; Thu, 7 Nov 91 23:36:29 +0100
Date: Thu, 7 Nov 91 23:36:29 +0100
From: dik@cwi.nl
Message-Id: <9111072236.AA15871@paring.cwi.nl>
To: mccalpin
Subject: streams results on 3090
Status: RO

The four files in the sharchive below are the stream results for the IBM 3090.
There are four files, either single or double precision and either non-
vectorized or vectorized.  The actual options used for the compiler are in
the first line of each file.  General information:

System:	IBM 3090-J 6VF (6 Vector processors, only one used here)
OS:	AIX
Compiler:fvs version 1.22 (doubtful, the compiler does not give a version
	number; 'what' lists 1.22 for /usr/bin/fvs)

The instructions told me to increase the number of array elements, I tried
multiplying by 10 but that appears to result in severe swapping (memory
utilization is erratic on this machine as it runs both AIX and VM/CMS so
you can not get the complete machine under AIX), so I did set the number
of elements to 500000 for each test.  The granularity of the clock is good
enough to get correct timing also for small times.

I will also do the streams test on the NEX SX-3 but he is currently not
accessible.  I noted in your results a few FPS results.  Are those results
of the new SPARC based FPS?  If that is not the case I can also run the set
on our FPS.

dik
--
dik t. winter, cwi, amsterdam, nederland
dik@cwi.nl
--
#! /bin/sh
# This is a shell archive, meaning:
# 1. Remove everything above the #! /bin/sh line.
# 2. Save the resulting text in a file.
# 3. Execute the file with /bin/sh (not csh) to create:
#	out.do
#	out.dov
#	out.so
#	out.sov
# This archive created: Thu Nov  7 23:27:50 1991
export PATH; PATH=/bin:/usr/bin:$PATH
echo shar: "extracting 'out.do'" '(714 characters)'
if test -f 'out.do'
then
	echo shar: "will not over-write existing file 'out.do'"
else
sed 's/^X//' << \SHAR_EOF > 'out.do'
XOptions: -f"optimize(3)"
X--------------------------------------
X Double precision appears to have 16 digits of accuracy
X Assuming 8 bytes per DOUBLEPRECISION word
X--------------------------------------
X Timing calibration ; time =   14.6361958235502243      hundredths of a second
X Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
X ---------------------------------------------------
XFunction     Rate (MB/s)  RMS time   Min time  Max time
XAssignment:  117.5385      0.0695      0.0681      0.0712
XScaling   :   95.4976      0.0842      0.0838      0.0854
XSumming   :  132.4796      0.0932      0.0906      0.0941
XSAXPYing  :  104.6561      0.1156      0.1147      0.1162
SHAR_EOF
if test 714 -ne "`wc -c < 'out.do'`"
then
	echo shar: "error transmitting 'out.do'" '(should have been 714 characters)'
fi
fi
echo shar: "extracting 'out.dov'" '(731 characters)'
if test -f 'out.dov'
then
	echo shar: "will not over-write existing file 'out.dov'"
else
sed 's/^X//' << \SHAR_EOF > 'out.dov'
XOptions: -f"optimize(3) vector(noreport)"
X--------------------------------------
X Double precision appears to have 16 digits of accuracy
X Assuming 8 bytes per DOUBLEPRECISION word
X--------------------------------------
X Timing calibration ; time =   7.89859667420387268      hundredths of a second
X Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
X ---------------------------------------------------
XFunction     Rate (MB/s)  RMS time   Min time  Max time
XAssignment:  251.9291      0.0324      0.0318      0.0331
XScaling   :  250.0782      0.0324      0.0320      0.0330
XSumming   :  288.5100      0.0453      0.0416      0.0472
XSAXPYing  :  264.5692      0.0458      0.0454      0.0465
SHAR_EOF
if test 731 -ne "`wc -c < 'out.dov'`"
then
	echo shar: "error transmitting 'out.dov'" '(should have been 731 characters)'
fi
fi
echo shar: "extracting 'out.so'" '(702 characters)'
if test -f 'out.so'
then
	echo shar: "will not over-write existing file 'out.so'"
else
sed 's/^X//' << \SHAR_EOF > 'out.so'
XOptions: -f"optimize(3)"
X--------------------------------------
X Single precision appears to have  6 digits of accuracy
X Assuming 4 bytes per default REAL word
X--------------------------------------
X Timing calibration ; time =   10.4539928      hundredths of a second
X Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
X ---------------------------------------------------
XFunction     Rate (MB/s)  RMS time   Min time  Max time
XAssignment:   92.8046      0.0550      0.0431      0.0566
XScaling   :   56.0577      0.0719      0.0714      0.0723
XSumming   :   72.5750      0.0832      0.0827      0.0839
XSAXPYing  :   68.5486      0.0971      0.0875      0.0995
SHAR_EOF
if test 702 -ne "`wc -c < 'out.so'`"
then
	echo shar: "error transmitting 'out.so'" '(should have been 702 characters)'
fi
fi
echo shar: "extracting 'out.sov'" '(719 characters)'
if test -f 'out.sov'
then
	echo shar: "will not over-write existing file 'out.sov'"
else
sed 's/^X//' << \SHAR_EOF > 'out.sov'
XOptions: -f"optimize(3) vector(noreport)"
X--------------------------------------
X Single precision appears to have  6 digits of accuracy
X Assuming 4 bytes per default REAL word
X--------------------------------------
X Timing calibration ; time =   5.22439957      hundredths of a second
X Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
X ---------------------------------------------------
XFunction     Rate (MB/s)  RMS time   Min time  Max time
XAssignment:  202.4803      0.0237      0.0198      0.0242
XScaling   :  166.4794      0.0242      0.0240      0.0243
XSumming   :  172.9955      0.0348      0.0347      0.0349
XSAXPYing  :  173.8526      0.0346      0.0345      0.0347
SHAR_EOF
if test 719 -ne "`wc -c < 'out.sov'`"
then
	echo shar: "error transmitting 'out.sov'" '(should have been 719 characters)'
fi
fi
exit 0
#	End of shell archive

From dik@cwi.nl  Fri Nov  8 08:01:02 1991
Received: from charon.cwi.nl by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA27310; Fri, 8 Nov 91 08:01:02 EST
Received: by charon.cwi.nl with SMTP; Fri, 8 Nov 1991 14:01:57 +0100
Received: by paring.cwi.nl ; Fri, 8 Nov 91 14:01:55 +0100
Date: Fri, 8 Nov 91 14:01:55 +0100
From: dik@cwi.nl
Message-Id: <9111081301.AA17636@paring.cwi.nl>
To: mccalpin
Subject: streams on NEC SX3
Status: RO

The two files in the sharchive below are the stream results for the NEC SX3.
General information:

System:	NEC SX3/14  single processor, four vector pipes
OS:	SX/UX
Compiler:f77sx Rev.012
Option:	-O

The instructions told me to increase the number of array elements.  The number
of elements in the timings was 5,000,000 (on the IBM 3090 in my previous
mail it was 500,000).  Still does not make it the initial calibration 30,
but also here the clock has adequate resolution.

Feel free to ask any questions you have,

dik
--
dik t. winter, cwi, amsterdam, nederland
dik@cwi.nl
--
#! /bin/sh
# This is a shell archive, meaning:
# 1. Remove everything above the #! /bin/sh line.
# 2. Save the resulting text in a file.
# 3. Execute the file with /bin/sh (not csh) to create:
#	out.single
#	out.double
# This archive created: Fri Nov  8 12:34:35 1991
export PATH; PATH=/bin:/usr/bin:$PATH
if test -f 'out.single'
then
	echo shar: "will not over-write existing file 'out.single'"
else
cat << \SHAR_EOF > 'out.single'
--------------------------------------
 Single precision appears to have  6 digits of accuracy
 Assuming 4 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 2.176048 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment: 4151.4609      0.0096      0.0096      0.0096  
Scaling   : 3835.5132      0.0104      0.0104      0.0104  
Summing   : 4993.7852      0.0120      0.0120      0.0120  
SAXPYing  : 4993.8984      0.0120      0.0120      0.0120  
SHAR_EOF
fi
if test -f 'out.double'
then
	echo shar: "will not over-write existing file 'out.double'"
else
cat << \SHAR_EOF > 'out.double'
--------------------------------------
 Double precision appears to have 16 digits of accuracy
 Assuming 8 bytes per DOUBLEPRECISION word
--------------------------------------
 Timing calibration ; time = 2.176072797738016 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function     Rate (MB/s)  RMS time   Min time  Max time
Assignment: 8302.1617      0.0096      0.0096      0.0096  
Scaling   : 7670.8556      0.0104      0.0104      0.0104  
Summing   : 9712.7843      0.0124      0.0124      0.0124  
SAXPYing  : 9712.9248      0.0124      0.0124      0.0124  
SHAR_EOF
fi
exit 0
#	End of shell archive

From cmg@magnet.cray.com  Mon Nov 11 21:46:23 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02090; Mon, 11 Nov 91 21:46:23 EST
Received: from sequoia.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6l)
	id AA25511; Mon, 11 Nov 91 20:47:30 CST
Received: from magnet (magnet.cray.com) by sequoia.cray.com
	id AA20408; 4.1/CRI-5.5; Mon, 11 Nov 91 20:47:30 CST
Received: by magnet (5.64/A/UX-2.00 Matt test 1b)
	id AA02759; Mon, 11 Nov 91 20:49:07 CST
From: cmg@magnet.cray.com (Charles Grassl)
Message-Id: <9111120249.AA02759@magnet>
Subject: stream
To: mccalpin
Date: Mon, 11 Nov 91 20:48:52 CST
X-Mailer: ELM [version 2.2 PL0]
Status: RO

Hello John;

I noticed the high SAXPY bandwidth for the CM-2 in your stream data
table.  Andrew Zackary brought this to my attention and also told me of
his correspondence with you.  

As each FPU in the CM-2 has one 32-bit wide path to memory, the peak
bandwidth would seem to be 7 MHz * 2000 CPUs * 4 Bytes, or 56,000
Mbyte/sec.  Do you know if the report 80593 Mbyte/sec is genuine or how
it was attained?

Also, I noted your entry for the NEC SX-3/14.  The SX-3/14 has a major clock
period of 5.8 nanoseconds (172 MHz) and a minor clock period of 2.9
nanoseconds (345 MHz).  The memory ports operate at 345 MHz in vector
mode.  

The SX-3/14 bandwidth seems rather low.  Two CPUs share three 256-bit
wide memory ports.  I do not know if these 256-bit wide ports go all
the way to shared memory, or if they go to another level of caching,
like in an IBM system.  Do you know much about the memory
architecture?

Regards,
-- 
Charles M. Grassl
Cray Research, Inc.
(612) 683-3531 cmg@cray.com

From cmg@magnet.cray.com  Tue Nov 12 09:25:31 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA02746; Tue, 12 Nov 91 09:25:31 EST
Received: from sequoia.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6l)
	id AA04015; Tue, 12 Nov 91 08:26:40 CST
Received: from magnet (magnet.cray.com) by sequoia.cray.com
	id AA26995; 4.1/CRI-5.5; Tue, 12 Nov 91 08:26:39 CST
Received: by magnet (5.64/A/UX-2.00 Matt test 1b)
	id AA03043; Tue, 12 Nov 91 08:28:22 CST
From: cmg@magnet.cray.com (Charles Grassl)
Message-Id: <9111121428.AA03043@magnet>
Subject: Re:  stream
To: mccalpin (John D. McCalpin)
Date: Tue, 12 Nov 91 8:28:14 CST
In-Reply-To: <9111121336.AA02579@perelandra.cms.udel.edu>; from "John D. McCalpin" at Nov 12, 91 8:36 am
X-Mailer: ELM [version 2.2 PL0]
Status: RO

Hello John;

FYI, some details regarding the NEC SX-3/14 architecture:

Each CPU has four independent double-pipes, for a total of eight
pipes.  A pipe is a multiple and and add functional unit.

Each of a pair of CPUs has has access to three memory ports:  two reads
and one write.  Each port is 256-bits, or four words, wide.  It is not
wide enough to support a SAXPY running on all eight pipes.  So, even
though the peak speed of the individual CPU is 5500 Mflop/s, on a SAXPY
its peak speed is 2750 Mflop/s.

A -VERY- ineresting test would be to run your stream program on two
CPUs of a SX-3.  It would be interesting to see how the shared memory
port contention is resolved.

The Fujitsu VP2000 has a similar shared hardware design.  In it, two
CPUs share the actual vector functional units and the memory ports.
The vector registers are duplicated for each CPU, but they also share
paths to memory.

As you have probably seen from some data for the CRAY Y-MP C90, each
CPU has a complete set of three 128-bit wide paths to memory.



Regarding the TMC results, could I look at the source code for the
test?  I think that I could get someone at the Minnesota Supercomputer
Center to run it on a CM-2.  As a matter of fact, I would like to try 
it myself.

Regards,
-- 
Charles M. Grassl
Cray Research, Inc.
(612) 683-3531 cmg@cray.com

From cmg@magnet.cray.com  Wed Nov 13 14:04:49 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA05178; Wed, 13 Nov 91 14:04:49 EST
Received: from sequoia.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6l)
	id AA03706; Wed, 13 Nov 91 13:05:53 CST
Received: from magnet (magnet.cray.com) by sequoia.cray.com
	id AA25607; 4.1/CRI-5.5; Wed, 13 Nov 91 13:05:51 CST
Received: by magnet (5.64/A/UX-2.00 Matt test 1b)
	id AA04251; Wed, 13 Nov 91 13:07:34 CST
From: cmg@magnet.cray.com (Charles Grassl)
Message-Id: <9111131907.AA04251@magnet>
Subject: stream
To: mccalpin
Date: Wed, 13 Nov 91 13:07:30 CST
X-Mailer: ELM [version 2.2 PL0]
Status: RO

Hello John;

I'm still trying to figure out how the CM-2 gets 80,000 Mbyte/sec.  I
suspect that the compiler has "optimized away" dead code.  Most current
compilers compilers should be able to delete some dead code is a test
program such as yours.

Below are results for 1, 2, 4, 8  and 16  CPUs on a CRAY Y-MP C90.
Hope that you can soon include these in your table.

Regards,
-- 
Charles M. Grassl
Cray Research, Inc.
(612) 683-3531 cmg@cray.com

1 CPU
--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 11.02508232 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function      Rate (MB/s)   RMS time    Min time   Max time
Assignment:  6965.4395       0.0368       0.0368       0.0368  
Scaling   :  6965.3798       0.0368       0.0368       0.0368  
Summing   :  9378.7332       0.0411       0.0409       0.0413  
SAXPYing  :  9500.7178       0.0404       0.0404       0.0404  



2 CPUs
--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 5.51878446 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function      Rate (MB/s)   RMS time    Min time   Max time
Assignment: 13865.9708       0.0185       0.0185       0.0185  
Scaling   : 13905.4840       0.0184       0.0184       0.0184  
Summing   : 18233.2027       0.0211       0.0211       0.0211  
SAXPYing  : 18246.2586       0.0211       0.0210       0.0211  



4 CPUs
--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 2.76657822 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function      Rate (MB/s)   RMS time    Min time   Max time
Assignment: 27610.2594       0.0093       0.0093       0.0093  
Scaling   : 27789.5522       0.0092       0.0092       0.0092  
Summing   : 34633.3071       0.0111       0.0111       0.0112  
SAXPYing  : 35044.1485       0.0110       0.0110       0.0110  



8 CPUs
--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 1.4161938 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function      Rate (MB/s)   RMS time    Min time   Max time
Assignment: 55071.9123       0.0047       0.0046       0.0048  
Scaling   : 55391.7676       0.0046       0.0046       0.0047  
Summing   : 60843.3497       0.0063       0.0063       0.0064  
SAXPYing  : 63229.5729       0.0061       0.0061       0.0063  



16 CPUs
--------------------------------------
 Single precision appears to have 14 digits of accuracy
 Assuming 8 bytes per default REAL word
--------------------------------------
 Timing calibration ; time = 0.82603542 hundredths of a second
 Increase the size of the arrays if this is <30  and your clock precision is =<1/100 second
 ---------------------------------------------------
Function      Rate (MB/s)   RMS time    Min time   Max time
Assignment:105497.3864       0.0024       0.0024       0.0024  
Scaling   :104654.3724       0.0025       0.0024       0.0025  
Summing   :101736.0623       0.0038       0.0038       0.0038  
SAXPYing  :103812.8177       0.0037       0.0037       0.0038  

From cmg@magnet.cray.com  Sun Nov 17 19:59:41 1991
Received: from timbuk.cray.com by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA10130; Sun, 17 Nov 91 19:59:41 EST
Received: from sequoia.cray.com by timbuk.cray.com (4.1/CRI-MX 1.6l)
	id AA04569; Sun, 17 Nov 91 19:01:09 CST
Received: from magnet (magnet.cray.com) by sequoia.cray.com
	id AA18265; 4.1/CRI-5.5; Sun, 17 Nov 91 19:01:08 CST
Received: by magnet (5.64/A/UX-2.00 Matt test 1b)
	id AA03065; Sun, 17 Nov 91 19:02:47 CST
From: cmg@magnet.cray.com (Charles Grassl)
Message-Id: <9111180102.AA03065@magnet>
Subject: Re:  stream
To: mccalpin (John D. McCalpin)
Date: Sun, 17 Nov 91 19:02:42 CST
In-Reply-To: <9111121439.AA02772@perelandra.cms.udel.edu>; from "John D. McCalpin" at Nov 12, 91 9:39 am
X-Mailer: ELM [version 2.2 PL0]
Status: RO

Hello John;

I'm still trying to figure out the CM-2 stream code.  I now strongly
suspect that the compiler has done some dead code elimination.

The orginal (from TMC) source code has additional "do i=1,100" loops
around each kernel and the TIMES array is adjusted accordingly.  I ran
the KAP output (from TMC) on a CRAY Y-MP (this code is attached
below).  The results are -low- by a factor of 100.  The KAP output does
not have the correction for the factor of 100.

I'm disappointed that the CF77 compiler did not see the dead
code in the "do i=1,100" loops.  Just the same, the calibaration
in this program is incorrect.

I suspect that the CM-5 compiler code optimized eventually kicked in
and deleted somethin.  Else, if it it really ran the program correctly,
the bandwidth is 8,000,000 Mbyte/sec.!

Regards,
-- 
Charles M. Grassl
Cray Research, Inc.
(612) 683-3531 cmg@cray.com

C Source code output from KAP:
      PROGRAM stream
C     .. Parameters ..
      INTEGER n,ntimes
      PARAMETER (p=256,n=4000*p,ntimes=10)
C     ..
C     .. Local Scalars ..
      INTEGER j,k,nbpw
C     ..
C     .. Local Arrays ..
      real a(n),b(n),c(n),maxtime(4),mintime(4),rmstime(4),
     $     times(4,ntimes)
      INTEGER bytes(4)
      CHARACTER label(4)*11
C     ..
C     .. External Functions ..
      INTEGER realsize
C     ..
C     .. Intrinsic Functions ..
      INTRINSIC dble,max,min,sqrt
C     ..
C     .. Data statements ..
      DATA rmstime/4*0.0/,mintime/4*1.0E+36/,maxtime/4*0.0/
      DATA label/' Assignment:',' Scaling   :',' Summing   :',
     $     ' SAXPYing  :'/
      DATA bytes/2,2,3,3/
      etime()=second()
C     ..
 
*       --- SETUP --- determine precision and check timing ---
 
      nbpw = realsize()
 
      t = etime()
          A = 1.0e0
          B = 2.0e0
          C = 0.0e0
      t=etime()-t
      PRINT *,'Timing calibration ; time = ',t*100,' hundredths',
     $  ' of a second'
      PRINT *,'Increase the size of the arrays if this is <30 ',
     $  ' and your clock precision is =<1/100 second'
      PRINT *,'---------------------------------------------------'
 
*       --- MAIN LOOP --- repeat test cases NTIMES times ---
      DO 60 k = 1,ntimes
 
         t=etime()
         DO 20 I=1,100
            C = A
 20      CONTINUE
         t=etime()-t
         times(1,k) = t
 
         t=etime()
         DO 30 I=1,100
            C = 3.0 * A
 30      CONTINUE
         t=etime()-t
         times(2,k) = t
 
         t=etime()
         DO 40 I=1,100
            C = A + B
 40      CONTINUE
         t=etime()-t
         times(3,k) = t

         t=etime()
         DO 50 I=1,100
            C = A + 3.0 * B
 50      CONTINUE
         t=etime()-t
         times(4,k) = t
	 call dummysub(a,b,c,n)
 60   CONTINUE
 
*       --- SUMMARY ---
C*$*NOVECTORIZE
      DO 80 k = 1,ntimes
         DO 70 j = 1,4
            rmstime(j) = rmstime(j) + times(j,k)**2
            mintime(j) = min(mintime(j),times(j,k))
            maxtime(j) = max(maxtime(j),times(j,k))
 70      CONTINUE
 80   CONTINUE
      WRITE (*,FMT=9000)
      DO 90 j = 1,4
         rmstime(j) = sqrt(rmstime(j)/float(ntimes))
         WRITE (*,FMT=9010) label(j),n*bytes(j)*nbpw/mintime(j)/1.0e6,
     $        rmstime(j),mintime(j),maxtime(j)
 90   CONTINUE
 
 9000 FORMAT (' Function',5x,'Rate (MB/s)  RMS time  Min time  Max time'
     $        )
 9010 FORMAT (a,4 (f10.4,2x))
      END
*-------------------------------------
* INTEGER FUNCTION realsize()
*
* A semi-portable way to determine the precision of default REAL
* in Fortran.
* Here used to guess how many bytes of storage a real number occupies.
*
	integer function realsize()
	double precision ref(30)
	real test
	double precision pi

C	Test #1 - compare double precision pi to acos(-1.0e0)

	pi = 3.14159 26535 89793 23846 26433 83279 50288 d0
	picalc = acos(-1.0e0)
	diff = abs(picalc-pi)
	if (diff.eq.0.0) then
	    print *,'Test #1 Failed = picalc=piexact'
	    print *,'Apparently Single=Double Precision'
	    print *,'Proceeding to Test #2'
	    print *,' '
	    goto 200
	else
	    ndigits = -log10(abs(diff))+0.5
	    goto 1000
	endif

C	Test #2 - compare single(1.0d0+delta) to 1.0e0

  200	do 10 j=1,30
	    ref(j) = 1.0d0+10.0d0**(-j)
   10	continue

	do 20 j=1,30
	    test = ref(j)
	    ndigits = j
	    call dummy(test,result)
	    if (test.eq.1.0e0) then
		goto 1000
	    endif
   20	continue
	print *,'Test #2 failed - Precision appears to exceed 30 digits'
	print *,'Proceeding to Test #3'
	goto 300

C	Test #3 - abs(sqrt(1.0d0)-sqrt(1.0e0))

  300	diff = abs(sqrt(1.0d0)-sqrt(1.0e0))
	if (diff.eq.0.0) then
	    print *,'Test Failed - sqrt(1.0e0)=sqrt(1.0d0)'
	    print *,'Apparently Single=Double Precision'
	    print *,'Giving up'
	    goto 400
	else
	    ndigits = -log10(abs(diff))+0.5
	    goto 1000
	endif


 1000	write (*,'(a)') '--------------------------------------'
	write (*,'(1x,a,i2,a)') 'Single precision appears to have ',
     $		ndigits,' digits of accuracy'
	if (ndigits.le.8) then
	    realsize = 4
	else 
	    realsize = 8
	endif
	write (*,'(1x,a,i1,a)') 'Assuming ',realsize,
     $                       ' bytes per default REAL word'
	write (*,'(a)') '--------------------------------------'
	return

  400	print *,'Hmmmm.  I am unable to determine the size of a REAL'
  	print *,'Please enter the number of Bytes per REAL number : '
  	read (*,*) realsize
	if (realsize.ne.4.and.realsize.ne.8) then
	    print *,'Your answer ',sizeof,' does not make sense!'
	    print *,'Try again!'
  	    print *,'Please enter the number of Bytes per ',
     $              'REAL number : '
  	    read (*,*) realsize
	endif
	print *,'You have manually entered a size of ',realize,
     $          ' bytes per REAL number'
	write (*,'(a)') '--------------------------------------'
	end

	subroutine dummy(q,r)
	r = cos(q)
	return
	end
        subroutine dummysub(a,b,c,n)
	return
	end

From dik@cwi.nl  Fri Nov 22 04:00:01 1991
Received: from charon.cwi.nl by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA08763; Fri, 22 Nov 91 04:00:01 EST
Received: by charon.cwi.nl with SMTP; Fri, 22 Nov 1991 10:01:49 +0100
Received: by paring.cwi.nl ; Fri, 22 Nov 91 10:01:47 +0100
Date: Fri, 22 Nov 91 10:01:47 +0100
From: dik@cwi.nl
Message-Id: <9111220901.AA00373@paring.cwi.nl>
To: mccalpin
Subject: SX3 benchmark
Status: RO

I have been informed that the SX3 I did use was not a SX3/14 but a SX3/12
(which means single processor, 2 vector pipes; not 4 vector pipes).
It appears that the memory bandwidth measures is indeed what NEC
representatives calculate.

dik

From sandee@Think.COM  Mon Dec 23 16:47:18 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA15536; Mon, 23 Dec 91 16:47:18 EST
Return-Path: <sandee@Think.COM>
Received: from Lugh.Think.COM by mail.think.com; Mon, 23 Dec 91 16:51:14 -0500
From: Daan Sandee <sandee@Think.COM>
Received: by lugh.think.com (4.1/Think-1.0C)
	id AA27728; Mon, 23 Dec 91 16:51:13 EST
Date: Mon, 23 Dec 91 16:51:13 EST
Message-Id: <9112232151.AA27728@lugh.think.com>
To: mccalpin
Subject: Re: Fortran 90 in the US DoD
Cc: sandee@Think.COM
Status: RO

|> I have a problem.  I am convinced that the memory bandwidth results
|> that alex@think.com sent me are bogus.  The optimizer is obviously
|> removing all the work and so just a timer latency is being reported.  
|> I am certain of this because Alex sent me the code that he used and the
|> MB/s rates do not include the factor of 100 that he put in to get a
|> timeable interval!  If his timings are correct, then the bandwidth is
|> actually 160 times the machine maximum, rather than just the 1.6 times
|> faster that he reports.

 I assume this is the same issue that I screwed up over in early November
or whenever. I remember you saying Alex reported bandwidth gt speedoflight then.
Obviously, a factor of 160 in excess, when he alleges a timer loop of 100,
smells like either the optimizer issue or just bad reporting ; i.e., he
didn't do the arithmetic right. Note on a MPP you can also screw up the
arithmetic easily be applying the number of processors wrong.
|> 
|> The code that I have to outsmart the optimizer probably will not work
|> very well on a CM-2, since it involves doing some scalar stuff to
|> selected array elements to confuse the optimizer.  

I never have much problem with the optimizer being too clever ...

|>          ...I can prepare a test
|> code that does "real" work, and so cannot be optimized, but I need a
|> good estimate of the precision and latency of the clock.  TMC would
|> probably also prefer that I run the test cases on a machine that is
|> less loaded than the CMNS machine -- the front ends there are pretty
|> busy, and will probably bias the results downward.

I did some mini-tests on a Sun4 of 25 MHz and a CM-2 at 8 MHz.
Results :
-  cm_timer_start plus cm_timer_stop take around 5 msec
    (this is presumably FE time)
-  for:
    do (n) times
     a=a+1
    enddo
   do loop overhead around 0.5 msec on unoptimized code,
                  around 7 microseconds on optimized code.
   This last is around the time you get out of the tables for a 64-bit add.

|> I would like to get this correction out soon.   The people at CRI
|> are (quite reasonably) upset at the impossible number that I posted,
|> and I would like to set the record straight....

Give me a simple code and I'll run it for you. Have to be tomorrow, though.
I'm in Tlh now, typing away at a Zenith with all workstations down. 
Email to this address as SCRI machines are down.

Daan Sandee                                           sandee@think.com
Thinking Machines Corporation
Cambridge, Mass 02142

From sandee@Think.COM  Tue Dec 24 14:06:39 1991
Received: from Mail.Think.COM by perelandra.cms.udel.edu (5.52/890607.SGI)
	(for mccalpin) id AA16187; Tue, 24 Dec 91 14:06:39 EST
Return-Path: <sandee@Think.COM>
Received: from Lugh.Think.COM by mail.think.com; Tue, 24 Dec 91 14:10:40 -0500
From: Daan Sandee <sandee@Think.COM>
Received: by lugh.think.com (4.1/Think-1.0C)
	id AA03096; Tue, 24 Dec 91 14:10:39 EST
Date: Tue, 24 Dec 91 14:10:39 EST
Message-Id: <9112241910.AA03096@lugh.think.com>
To: mccalpin
Subject: Re: CM2 bandwidth
Status: RO

Your program was using cm_timer_cm_read_busy, presumably inserted by Alex.
Unfortunately, this function and other functions that return a CM time to
the caller are BROKE on CMSS 6.1. So I had to revert to the usual
CM_timer_print which prints to standard error.
Also, the optimizer did optimize away all of your code.
So I returned to my own program, and did basically the same thing, using
primitives that were *not* optimized away.
Now I can't give you a spread with RMS error on the times, but just by
looking at consecutive runs I can tell you that the number of digits I
supply indicates the accuracy.
The following are for N=1,000,000  repeat count NITER=125, making for
exactly 1 GB (decimal, that is) per stream (double precision).
Times in seconds.

                  CM-2 8K 8MHz   CM-2 4K 8MHz    CM-2 4K 7MHz   bytes/tick
 *optimized*        (256 PEs)      (128 PEs)       (128 PEs)    (approx)
  a = a + 1.d0        0.523          1.045           1.201           2
  a = a + b           0.713          1.425           1.633           2
  a = a + 2.d0*b      0.717          1.434           1.642           2
 *unoptimized*
  a = 2.d0            0.364          0.727           0.833           1.333
  a = b               0.543          1.085           1.242           2
  a = 2.d0*a          0.523          1.046           1.201           2
  a = b + c           0.716          1.432           1.642           2
  a = c + 2.d0*b      0.750          1.500           1.719           2

This looks like a speed of 2 bytes per PE per clock tick.
Which is less than the nominal rate of 1 slice per PE per tick.

Daan Sandee                                           sandee@think.com
Thinking Machines Corporation
Cambridge, Mass 02142

