************************************************************* 2.24.91 * STACK.TALK * BY * EV VERGUIZAS * ************************************************************* HOW TO GET CERTAIN STACK-BASED SCREEN UPDATES TO RUN FASTER. If you are new to this topic, stop right here. You should read DYA Jim Maricondo's TechNote #3, Jonah Stich's source code for MegaScroll II, and Apple's (GS) Technote #70 before tackling this. See also Tilescroll by Kevin Grossnicklaus. All of these are available on AO. Per Jim, to get best results in using the stack with the PEI instruc- tion for graphics updates you must keep the Direct Page Register (DPR) set to $xx00. This has the effect of reducing the machine cycles for each PEI by one. It is fairly simple to keep the DPR at $xx00 when you are copying a continuous field of bytes. (See Jim's article on this.) When you have to keep the DPR's low byte at zero for a non-continuous field of bytes, such as a "viewscreen", it is not so obvious how to do this. This algorithm is directly applicable to situations in which 1) you have a field of bytes which does not go edge-to-edge on screen (that is, you have the typical inset viewscreen a la Tunnels, or MegaScroll, or Tile- scroll), 2) you have first written the entire field of bytes to shadow screen with shadowing off, and 3) then you turn on both the machine state and shadow switches and copy everything on the viewscreen to itself using the stack. What follows is an apparently complex technique for keeping the low byte of the DPR at $00. I say apparently because, as you will see, behind the superficial complexity lies a nice symmetry. I've used for an example here the coordinates and screen size in Kevin's Tilescroll demo. The example is specific to Kevin's code but the method is not. What follows immediately is the METHOD for determining where the Stack ptr shall point, to what value the DPR shall be set, and what exact PEI's must be used. In the following discussion refer to the chart entitled "Pattern". First, find the original value of your stack ptr, i.e, normally, the bottom right hand corner where shadowing is to start (here, 31201). Now, by subtracting 160 successively, determine the value of each of the next seven starting points of the lines directly above the original position. (See Stack Ptr column in Pattern.) Second, find the value of the DPR which ends in $00 which is just less than the value of the original stack ptr (here 30976 (=$7900) is the first value ending in $00 which is less than 31201). Now, subtract successive- ly 256 ($100) from 30976 until you have seven. (See DPR column in Pattern.) (In practice, you will not need seven. This is for the exercise.) Line 1. We line up our stack ptr and DPR's on opposite sides of the paper and figure out which stack ptrs belong to which DPR and what PEI's to use. (Bear in mind that the DP has a reach of only 256 bytes so the stack ptr must be pointing to a byte no farther up in memory than 256 b. from the DPR.) DPR 30976 is only 226 bytes down (in memory) from 31201. So we need to start with the 113th PEI instruction (rem: each PEI lays down 2 bytes). No. 113 overlays the byte pointed to by the stack ptr and, as the stack automatically decrements, the decreasing PEI's keep perfect pace. No. 113 is PEI $e0. So, since the width of our viewscreen is 100 bytes (b.), we need to PEI down 100 b. (50 PEI instructions),i.e., through and including PEI $7e. Line 2. The next line (we just copied one) has to start at 31041. (We'll get into how to do this later.) At this point, looking at the current DPR, we see it is still 30976, as that continues to be less than the new stack ptr. (We cannot use the next lower DPR, 30720 (=$7800), because 30720 is farther than 256 bytes down in memory from 31041.) We now calculate how many bytes we can copy using the current DPR before it is used up and we have to switch to the next lower one. We now calculate: 31042 (that's right) - 30976= 66 b. This means that we count up 33 PEI instructions from, and including, PEI $00, and this puts us at PEI $40. So the code will go PEI $40 thru PEI $00. Line 2 (cont.). At this point we have shadowed 66 of the 100 bytes needed for line 2. We need another 34 to finish, but we have used up the current DPR and must now switch to the next lower. The situation is illustrated here: NEXT CURRENT LOWER DPR= DPR= $7900 $7800 ________________________________________________________________________________ DP OFF- | SET IN | BYTES: $00 $01................ $FA $FB $FC $FD $FE $FF | $00 $01 $02 $03 _ ------- ------- ------- INSTR'S LAYING DOWN BYTES ABOVE: PEI $FA PEI $FC PEI $FE PEI $00 PEI $02 If we count down from PEI $FE, the 17th instruction (34th b) is PEI $DE. Now, by PEI'ing from $FE through $DE we shadow the 34 b needed to complete the line. Line 3. To start line 3 the stack ptr has to be set to 30881. The DPR set to 30720 can be used for this (30882-30720=162 b) This means two things: this DPR is within 256 b of the stack ptr and so is usable, and, because you only have to copy 100 b, this particular line can be copied in its entirety without switching DPR's again. The 100 b are copied using PEI $FE through PEI $A0 (50 words= 100 b). Now that you understand the general procedure, carry on, in your own code, for about 10 or 12 lines, until you begin to see THE PATTERN. (This is only an exercise. When you actually code, you would not need to do this many lines.) What pattern, you say? If you look at the successive series of PEI's in the chart, Pattern, you will that the PEI's in line #1 are repeated in line 9, those of line 2 in 10,of 3 in 11, and so on. PATTERN LINE # DPR WHICH PEI'S TO USE AND #OF BYTES COPIED STACK POINTER 1 30976 PEI $E0 -> PEI $7E 100 BYTES 31201 2 30976 PEI $40 -> PEI $00 66 BYTES 31041 30720 PEI $FE -> PEI $DE 34 BYTES 3 30720 PEI $A0 -> PEI $3E 100 BYTES 30881 4 30720 PEI $00 2 BYTES 30721 30464 PEI $FE -> PEI $9E 98 BYTES 5 30464 PEI $60 -> PEI $00 98 BYTES 30561 30208 PEI $FE 2 BYTES 6 30208 PEI $C0 -> PEI $5E 100 BYTES 30401 7 30208 PEI $20 -> PEI $00 34 BYTES 30241 29952 PEI $FE -> PEI $BE 66 BYTES 8 29952 PEI $80 -> PEI $1E 100 BYTES 30081 9 29696 PEI $E0 -> PEI $7E 100 BYTES 29921 10 29696 PEI $40 -> PEI $00 66 BYTES 29761 29440 PEI $FE -> PEI $DE 34 BYTES 11 29940 PEI $A0 -> PEI $3E 100 BYTES 29601 12 29440 PEI $00 2 BYTES 29441 29184 PEI $FE -> PEI $9E 98 BYTES Kowabunga! If the screen can be shadowed using the stack by hard- coding the PATTERN of PEI's for lines 1-8, then, by changing the appropriate stack ptrs and setting the DPR's correctly, the same code can be reused for the next 8 lines, etc. Hmmm. We could then use a precalculated table for the DPR and Stack Ptr values. And then, so as not to have to increment the index every time a PEI series is done, have a different "Stack" table for the beginning of each line and a different DPR table for use once each DPR is "used up". Further hmmm. Also, the string of PEI's we use is really, in the case of stack-based, nothing but a measure of distance on screen. So, if the PEI series repeat, so should the DPR's and Stack Ptrs, but with an offset. Curiously, the DPR and Stack Ptr as set just before copying line 9 are each 1280 less than the original DPR and the original StackPtr. Why 1280? Be- cause it is the Least Common Multiple of the 256, 512,.. 1280 and 160,320,480... 1280 series. These two series get back to their original relation at offsets of -1280, -2560, etc. And, in between, line 10's DPR and Stack Ptr are offset 1280 from line 2's, line 11's 1280 from line 3's, etc. Every eight lines the DPR and the Stack Ptr are again the same "distance" they were from each other eight lines before. Kowabunga again! This makes coding the data areas a piece of cake: DP1 30976,30976-1280,30976-2*1280,... Stack1 31201,31201-1280,31201-2*1280... Now, here's how it all looks coded out: ldx #0 init loop loop lda stack1,x tcs lda dp1,x tcd pei $e0 ; .... all intervening PEI's pei $7e ;you don't need to inx inx here be- ;cause you have hardcoded the offsets ;for the first 8 lines lda stack2,x tcs ;you don't need to change the DPR yet ;the first 66 bytes of line 2 are sha- ;dowed with the old DPR (see above) pei $40 ; .... all intervening PEI's pei $00 lda dp2,x since the DPR changes for the last 34 tcd bytes of L. 2 you get the new DPR from pei $fe ;the table. Stack has decrem. on its own. ; .... Only time SPtr needs resetting is at pei $de beginning of each line. ..and so it goes, until the last segment: lda stack8,x the eighth stack ptr tcs lda dp5,x the 5th dpr tcd pei $80 ; ... all intervening PEI's pei $1e inx inx cpx #--- fill in 2 * no. of times this loop ; has to be repeated. (Depends on size of ; your viewscreen.) bcs rest_of_my_program brl loop else, do it again Examples of Data areas: stack1 anop dc I2'31201,31201-1280,31201-2*1280....' stack2 anop dc I2'31041,31041-1280,31041-2*1280....' dp1 anop dc I2'30976,30976-1280,30976-2*1280....' Well, there it is. This method works: it makes the shadowing subroutine run 14% faster than the comparable shadowing code using an uncontrolled DPR. It is reusable with no changes so long as you have a screen exactly the same size and at the same coordinates. Even if the screen is different the only real hassle is figuring out the first set of lines. As you have seen, the rest of the stuff is boilerplate consisting of various offsets. The disadvantages are that it takes up a lot more code and that it is very easy to get the various values wrong in your lines. If you follow the method above, however, you'll see that it soon becomes a lot easier. Now, two things to end this for now. First, because the offsets with- in each set of 8 lines are fixed for any given rectangular inset viewscreen, you can save a lot of code by simply using only the sec and sbc instructions to produce the appropriate offsets within the loop prior to each PEI-series. This costs some efficiency but saves lots of data. In this case, you would need only two data areas. The reason the code here is the longer version is that I wanted to show the fastest possible whole-screen update. Second, keep the height of your viewscreen a multiple of eight: that way you will never have to hardcode more than eight lines. Lastly, remember that the actual on-screen improvement in perfor- mance results not just from the shadowing code but from all of the code ne- cessary to get your images on screen. If you have been using an uncontrolled DPR in shadowing your whole viewscreen, this will improve the speed of your screen update. I am interested in comments on this technique. Has it been obviated by something? Is there a better workaround? I have only put this out because I have not seen it elsewhere and because, when I figured it out, I thought it was worth sharing. I check in at AO occasionally and can be reached by US Mail: Ev Verguizas 4478 S.W. 13th Terrace Miami, Fla. 33134. P.S. Many thanks to Parik Rao, Jonah Stich, Jim Maricondo, Chris McKinsey, Stephen Lepisto, and Kevin Grossnicklaus for sharing your knowledge.