Analysis model: gpt-5.5 xhigh
Xtal by Complex Media Labs - Technical Dissection
Scope
This is a binary and complete direct-runtime dissection of
xtal - the demonstration by Finnish group Complex Media Labs, released in
1995. The public production metadata identifies it as a PC demo, and the
release text inside the archive credits several important runtime components:
- Pouet: https://www.pouet.net/prod.php?which=71
- Demozoo production: https://demozoo.org/productions/28488/
- Demozoo Complex group page: https://demozoo.org/groups/175/
- Archive: https://files.scene.org/get/demos/groups/komplex/clx_xtal.zip
- PMODE/W v1.21 by Charles Scheffold and Thomas Pytel
- MIDAS v0.50 pre-release #1 by Sahara Surfers
- Runtime sine tables using Karl's sine generator
320x200x256 50hz modereverse-engineered by Liket
The caveat is important: this is not source code. The executable is a packed PMODE/W protected-mode binary. I flattened the PMODE/W image, inspected the recovered flat memory map, and ran the authenticated executable through its complete native no-sound route. A few PMODE/W relocation warnings were emitted while flattening, so control-flow and local loops are reliable enough for analysis, but any claim about high-level symbol names should be read as an inference from code shape, constants, data flow, and direct runtime concordance.
Examined Files
The archive contains:
XTAL.EXE PMODE/W-packed DOS executable
XTAL.XM FastTracker II module, title "gates of ixtlan"
XTAL.TXT release notes
file_id.diz BBS description
Hashes from the examined files:
349571136dd60871d9c4af32b278cedd2a46bb49879f3d8566f59cc6463d49f6 clx_xtal.zip
a31e3d1ec52d018158935c2665f86701b36da54ac322f629a61d09f06f8dc87d XTAL.EXE
b9163e835e645edb33b6bc73b57c398ac8e06c137c798ef07e7212fc863dd9ad XTAL.XM
XTAL.TXT gives the most useful non-code hints:
- doesnt work with win95 :)
- stereo mode is forced to mono, except on gus...
- music player is midas v0.50 pre-release #1 by sahara surfers
- dos extender is pmode/w v1.21 by charles scheffold and thomas pytel
- sine tables are runtime calculated using karl's sine generator
- 320x200x256 50hz mode is reverse-engineered by liket
file_id.diz describes the release as the final 4 MB memory version.
Complete Direct Silent Runtime
The earlier runtime pass used DOSBox-X's internal H.264 recorder, selected Gravis UltraSound, stopped on a schedule after 234.6 seconds, and then killed DOSBox. It showed useful imagery but did not prove the endpoint.
The new run uses the authenticated archive and MIDAS's own No Sound device.
No guest sound card is selected, SDL uses a dummy host driver, and the external
FFV1 master contains video only. The visual scheduler advances normally and
the executable returns naturally after its final credit card.
Emulator: DOSBox-X 2026.01.02
Machine: S3 VGA, 16 MB RAM
CPU: normal 486 core, fixed 50,000 cycles
Memory: XMS on; EMS and UMB off
Show path: original XTAL.EXE, MIDAS "No Sound"
Capture: external X11 grab, 25 fps FFV1, no audio stream
Working master: 226.2 seconds including setup
Endpoint: natural return to DOS; DOSBox exit status 0
Published media: six continuous GIF excerpts; no audio or H.264/MP4
Approximate direct-capture phases:
| Capture time | Directly observed state |
|---|---|
00:00–00:11 |
MIDAS setup and the selected No Sound path. |
00:12–00:32 |
Handwritten Complex logo, production card, and spaced xtal title. |
00:33–00:52 |
Magenta and green radial bursts establish the shared tunnel field. |
00:53–01:13 |
Painted face and luminous wire/mesh layers blend into the tunnel. |
01:14–01:43 |
Face, eye and concentric material repeatedly enter the same radial transform. |
01:44–02:43 |
Bright starbursts, red rings and dense high-contrast tunnel phases. |
02:44–03:26 |
Green, grey and white contour tunnels become increasingly warped and dense. |
03:27–03:46 |
Final tunnel dissolves into the complex / xtal / jmagic - jugi - reward card, then the executable returns to DOS. |
These are continuous excerpts from that single direct silent run:






Executable Shape
The MZ header describes a small real-mode loader followed by a PMODE/W payload:
XTAL.EXE size: 1141351 bytes
MZ image size: 9808 bytes
header paragraphs: 4
relocations: 0
initial CS:IP: 0000:0059
initial SS:SP: 0344:0100
The string PMODE/W v1.21 occurs in the real-mode stub, and the PMODE/W packed
payload marker PMW1 appears at file offset 0x2650, immediately after the MZ
loader image.
Flattening the PMODE/W image produced:
Flat memory map: /tmp/xtal-analysis/unpacked/XTAL.EXE.FLAT
Flat size: 3782003 bytes
Entry point: 0x003909D8
Stack pointer: 0x00012A90
The flat image contains ordinary Watcom runtime text near the entry point:
WATCOM C/C++32 Run-Time system.
That is consistent with a 32-bit protected-mode C or C++ program using runtime support code, direct hardware routines, and hand-written assembly inner loops.
High-Level Program Flow
At 0x3909d8, execution jumps over the Watcom copyright string into startup
code. The early path is mostly runtime and DOS-extender setup: stack, argument
environment, selectors, file and memory setup, then device and music setup.
The demo-specific setup path calls these important routines:
0x38a31a custom VGA mode setup
0x390646 retrace/PIT calibration
0x390739 IRQ0 vector installation path
0x11288f sine-table recurrence builder
0x1128ab shade/ramp table builder
The archive also stores the music as a separate XTAL.XM file. The executable
contains strings such as:
xtal.xm
MIDAS Error
Using %s
Extended Module
ULTRASND
BLASTER
So the binary loads the external XM module through MIDAS, probes common DOS audio configuration paths, and uses the music player's reported position as part of its effect scheduler.
The central scheduler repeatedly does three things:
- Run the current effect update or renderer.
- Poll the keyboard with BIOS
int 16h, ah=1and branch to exit if needed. - Call the timing/music-position updater at
0x1100a4and compare the resulting clock against phase thresholds.
The common phase value used by many sections is built from fields near
0x117dc..0x117e4. A typical compare loads:
movzx eax, byte [0x117dc]
shl eax, 8
or al, [0x117e4]
cmp ax, phase_limit
That makes the demo timeline music-driven rather than just frame-count driven.
The 50 Hz VGA Mode
The most explicit hardware trick is the custom 320x200x256 50 Hz VGA mode at
0x38a31a. It starts from BIOS mode 13h and then rewrites VGA timing registers:
mov ax, 0013h
int 10h
Then it unlocks protected CRTC registers:
mov dx, 03d4h
mov al, 11h
out dx, al
inc dl
in al, dx
and al, 7fh
out dx, al
CRTC register 0x11 is the vertical retrace end register. Bit 7 is the CRTC
write-protect bit. Clearing it allows writes to the vertical timing registers.
Next the code writes the VGA miscellaneous output register:
mov dl, 0c2h
mov al, 0e3h
out dx, al
Then it returns to the color CRTC index port and writes a full timing block.
The code uses 16-bit indexed VGA writes: AL is the register index, AH is the
value, and out dx, ax writes index to 0x3d4 and value to 0x3d5.
CRTC 00 = 60
CRTC 01 = 4f
CRTC 02 = 50
CRTC 03 = 82
CRTC 04 = 54
CRTC 05 = 80
CRTC 06 = 6f
CRTC 07 = 3e
CRTC 08 = 00
CRTC 09 = 41
CRTC 10 = f6
CRTC 11 = 88
CRTC 12 = 8f
CRTC 13 = 28
CRTC 14 = 40
CRTC 15 = 90
CRTC 16 = 6c
CRTC 17 = a3
The key consequences:
CRTC 13 = 0x28is the familiar mode-13h style scanline offset. In chained 256-color addressing this keeps the framebuffer as a simple 320 byte line surface.CRTC 06 = 0x6fplus overflow bits fromCRTC 07 = 0x3edecodes to a vertical total of about 625 scanlines. That is the PAL-like 50 Hz part.CRTC 12 = 0x8fplus overflow bits gives a displayed vertical end of about 400 hardware scanlines, matching the usual double-scanned 200-line VGA image.CRTC 10 = 0xf6places vertical retrace start late in the 625-line frame.CRTC 15 = 0x90starts vertical blanking around the end of visible display.
The horizontal setup is also not a vanilla BIOS table:
- Horizontal total register
0x00 = 0x60. - Horizontal display end
0x01 = 0x4f, which is 80 character clocks, or 320 visible pixels in 4-pixel 256-color shift terms. - Horizontal retrace begins at
0x04 = 0x54.
After CRTC setup, the code programs the sequencer:
SEQ 1 = 01
SEQ 3 = 00
SEQ 4 = 0e
SEQ 4 = 0x0e enables the mode-13h style memory behavior: extended memory,
odd/even disabled, and chain-4 addressing.
Then it programs the graphics controller:
GC 5 = 40
GC 6 = 05
GC 5 = 0x40 selects the 256-color shift behavior, and GC 6 = 0x05 selects
graphics mode with the A0000 aperture. The result is deliberately practical:
the video memory is still easy to address as a linear 64000-byte surface, but
the monitor timing has been bent to 50 Hz.
That is why the release note says the mode was reverse-engineered and does not work under Windows 95. This code assumes direct VGA I/O, a real retrace status bit, and hardware that accepts non-BIOS timing values.
Scanline Height Trick
The demo changes CRTC register 0x09, the maximum scanline register, in later
parts:
; around 0x1111d6
mov dx, 03d4h
mov ax, 4309h
out dx, ax
; around 0x112850
mov dx, 03d4h
mov ax, 4109h
out dx, ax
The normal custom mode sets CRTC 09 = 0x41; this later code temporarily uses
0x43. The low bits of this register control scanlines per character row. In a
256-color chained mode this changes the vertical stretching and row cadence
without changing the byte layout of the software buffers. That is a classic
VGA-screen trick: one effect can be made taller, chunkier, or scroller-like by
changing hardware row repetition instead of resampling the whole image.
Retrace And Timer Synchronization
The binary installs its own IRQ0 handler. The vector-management path around
0x390739 uses DOS interrupt vector calls:
mov ax, 3508h
int 21h ; get old IRQ0 vector
mov edx, 390392h
mov ax, 2508h
int 21h ; install new IRQ0 handler
The handler at 0x390392 saves all general registers and segment registers,
sets its data segments, and then branches on a state byte around 0xfc2e.
The interesting mode waits on the VGA input-status register:
mov dx, 03dah
wait_until_not_in_retrace:
in al, dx
test al, 08h
jne wait_until_not_in_retrace
; optional callback through [0xfc16]
; fixed-point counters updated here
wait_until_in_retrace:
in al, dx
test al, 08h
je wait_until_in_retrace
; optional callback through [0xfc1a]
call 3902ddh
mov al, 20h
out 20h, al ; PIC end-of-interrupt
; optional callback through [0xfc1e]
So IRQ0 is not just a plain timer tick. It is interlocked with the VGA vertical retrace bit. That lets the demo run frame callbacks at stable display moments, update counters, and avoid tearing-sensitive updates happening in the middle of visible scanout.
The calibration routine at 0x390646 measures the retrace period using the PIT:
; wait for a retrace transition through port 03dah
mov al, 36h
out 43h, al ; PIT channel 0, mode 3, lobyte/hibyte
xor al, al
out 40h, al
out 40h, al ; reload counter with 0
; wait for another retrace transition
mov al, 00h
out 43h, al ; latch counter 0
in al, 40h
mov ah, al
in al, 40h
xchg al, ah
neg ax ; elapsed count
It repeats the measurement and accepts it only when two readings differ by at most two PIT ticks. That is a good tell that the authors wanted a stable frame duration, not a rough delay.
Palette And DAC Paths
Xtal spends a lot of work on palette control. The direct DAC upload routine at
0x1100f8 writes a whole 256-color palette:
mov edx, 03c8h
xor al, al
out dx, al ; DAC write index = 0
mov ebx, palette
lea ecx, [ebx+0300h]
mov esi, 03c9h
upload:
mov al, [ebx+0]
mov edx, esi
out dx, al
mov al, [ebx+1]
out dx, al
mov al, [ebx+2]
add ebx, 3
out dx, al
cmp ebx, ecx
jne upload
That is exactly 0x300 bytes: 256 RGB triples. It sets the start index to zero
with port 0x3c8, then writes RGB data to port 0x3c9.
There is a second compact DAC upload at 0x111340:
mov dx, 03c8h
xor al, al
out dx, al
inc dl ; 03c9h
mov ecx, 0300h
rep outsb
Two special routines force the whole DAC to white or black:
; whiteout around 0x112798
mov dx, 03c8h
xor al, al
out dx, al
inc dl
not al ; ffh
mov ecx, 0300h
rep outsb-like loop
; blackout around 0x11287d
mov dx, 03c8h
xor al, al
out dx, al
inc dl
mov ecx, 0300h
rep zero output
Classic VGA DACs store 6-bit color values, so values above 0x3f are
effectively saturated or masked by the hardware. The intent is still clear:
instant full-screen white and black transitions without touching the framebuffer.
Resource Expansion
The image/resource unpacker at 0x1128ef is a tiny RLE decoder. It receives a
destination in EDI and an output byte count in ECX, then computes the end
pointer:
lea edx, [edi+ecx]
The stream is controlled by one byte at a time:
next_packet:
mov cl, [esi]
inc esi
shr cl, 1
jc run_packet
literal_packet:
mov al, cl
stosb
cmp edi, edx
jb next_packet
ret
run_packet:
mov al, [esi]
inc esi
rep stosb
cmp edi, edx
jb next_packet
ret
Odd control bytes mean "repeat the following byte control >> 1 times".
Even control bytes mean "emit control >> 1 as a literal byte". This is small
and fast, and it fits palette-index images well because the literal path does
not need to carry a second source byte.
The unpacker is used for both full-screen and off-screen buffers. One path
expands directly to 0xa0000, and others expand into software buffers that are
later transformed or copied.
Runtime Sine And Shade Tables
The release note about Karl's sine generator matches the code. Routine
0x11288f builds a recurrence table at runtime:
mov edi, 11291ch
mov ecx, 07feh
loop:
mov ebx, [edi-4]
mov eax, ebx
imul ebx
shrd eax, edx, 1dh
sub eax, [edi-8]
stosd
loop loop
This is not calling sin() from a library. It is using previous table values to
generate the next ones. The code then uses masked phase indexes such as
phase & 0x1ffc, and a +0x800 phase offset for the perpendicular component.
That gives sine/cosine pairs from one table.
Routine 0x1128ab builds a shade or distance ramp table near 0x12c144. Its
outer loops walk signed byte-like coordinates, and its inner part computes and
clamps a brightness:
; simplified shape
for bh from 7fh down through signed range:
for bl from 7fh down through signed range:
value = ecx + edx + 3fh
if high_byte(value) > 78h:
value = 7800h
store high_byte(value)
The exact table consumers show that this is used as a mapping table during texture distortion. A texture sample is not written directly; it is combined with a computed index and then looked up through this precomputed ramp. That keeps the inner loop byte-sized and avoids multiplies or square roots per pixel.
Palette Ramp Builder
The helper at 0x1107ef writes a 64-entry linear ramp. It takes a starting
fixed-point value and an increment, clamps the increment to 0..0x100, then
stores the high byte for 64 entries:
cmp edx, 100h
jbe increment_ok
mov edx, 100h
increment_ok:
test eax, eax
js negative_start
mov ecx, 40h
positive_loop:
stosb_as_ah
add eax, edx
loop positive_loop
The negative-start path emits zero until the accumulator becomes positive, then continues with the same high-byte ramp. This is exactly the kind of primitive you want for palette fades and color-table morphs: it turns a phase and slope into a compact 64-byte channel curve.
Higher-level routines combine these curves:
0x110828 builds a 16x16-ish table from channel ramps and a bias
0x1109e8 combines two 0x300-byte palette/table blocks into 0x114d0
0x110a20 drives palette/table morphing from phase variables
The important point is that palette motion is not just "fade current DAC toward black". The demo builds intermediate tables and uses them in render loops. That is why the color changes can be tied to distortion and shading, not only to global brightness.
Distortion Offset Table
Routine 0x110b6b builds a 256-entry word table at 0x110329. It reads the
runtime sine table, uses a small function table around 0x110bb0, and writes
per-line or per-column offsets:
phase = [0x1102a8] + frame_dependent_terms
for i in 0..255:
sample = sine_table[(phase + i * step) & mask]
offset_table[i] = transformed(sample)
Several variables around 0x1102a8, 0x1102b0, and related addresses feed
this builder. These are not static lookup tables baked into the executable; the
waves are recomputed as phase accumulators move.
This table is then consumed by the texture distortion path. The pattern is classic for mid-1990s VGA demo effects:
- Update wave phases once per frame.
- Build a small offset table.
- Use that table inside a byte-per-pixel loop to avoid expensive math.
Distorted Texture Sampler
Routine 0x110bf1 is one of the important visible-effect loops. It reads an
index table, adds a wave displacement, samples a texture, maps it through a
shade table, and writes an off-screen byte buffer.
The setup is roughly:
edi = 124434h ; destination buffer
ecx = height * 0a0h - 1
The hot path is:
loop:
; fetch base index from a word table
movzx eax, word [index_table + ...]
; add dynamic wave displacement
add ax, [110329h + ...]
; sample texture
movzx eax, byte [eax + 17550dh]
; shade/remap sample
mov al, [eax + 12c144h]
; write software framebuffer byte
stosb
dec ecx
jnz loop
This loop is important because it separates the expensive-looking visual from the real work:
- The geometric distortion is reduced to table additions.
- The texture is a byte-indexed source image.
- The lighting or crystal brightness is a lookup table.
- The destination is a linear byte buffer, not VGA planes.
That is why the effect can run on mid-1990s PC hardware. The inner loop does no division, no trigonometry, and no per-pixel port I/O.
Symmetric Texture Mapper
Routine 0x110cf3 is a denser mapper. It renders two pixels per iteration from
a source texture around 0x143e57, using fixed-point u/v accumulators and
then draws a mirrored half.
The hot two-pixel part looks like this:
mov eax, 004f0000h
span_loop:
mov ebx, ecx
shr ebx, 24
add ecx, esi
shld ebx, edx, 8
add edi, 2
mov al, [ebx + 143e57h]
add edx, ebp
mov ebx, ecx
shr ebx, 24
add ecx, esi
shld ebx, edx, 8
mov ah, [ebx + 143e57h]
add edx, ebp
mov [edi-2], ax
sub eax, 00010000h
jae span_loop
What this does:
ECXandEDXare fixed-point texture coordinates or coordinate terms.ESIandEBPare the per-pixel increments.shr ebx, 24extracts the high byte from one accumulator.shld ebx, edx, 8injects bits from the other accumulator, forming a compact 2D texture address.ALandAHreceive two adjacent output pixels.- A single 16-bit store writes both pixels.
After the forward span, the routine reverses direction for the mirrored half by subtracting the accumulators and writing backward. That matters: the effect is not rendering every pixel independently. It is using symmetry to halve setup work and keep the span loop tight.
The texture address construction is also a neat trick. Instead of computing
y * width + x, the texture is arranged so that selected high bits from two
fixed-point accumulators can be fused into a usable address. That is a very
demo-scene way to trade memory layout for speed.
Nibble Compression And VRAM Expansion
The software effect buffers are not always stored in final VGA form. Routine
0x110c3b processes 0x1f40 dwords with a mask:
mov edx, 0f0f0f0fh
mov ecx, 1f40h
loop:
mov eax, [esi]
and eax, edx
; optional shift by 4 depending on [0x1102f8]
mov [esi], eax
add esi, 4
loop loop
0x1f40 dwords is 32000 bytes. The 0x0f0f0f0f mask keeps low nibbles from
each byte. The optional shift selects the other nibble phase. This is a cheap
way to prepare a half-resolution, 4-bit, or interleaved-looking buffer before
the final VGA write.
Routine 0x110c7d expands intermediate data into actual VGA memory near
0xa0000. The loop uses a base table around 0x114a30, reads packed nibbles
from 0x124434, and writes dwords to the framebuffer:
esi = 1
height = [0x1107eb]
row_loop:
ecx = 50h ; 80 dwords = 320 bytes
column_loop:
; unpack or map two nibbles through a table
; build one 32-bit group of VGA palette indexes
mov [esi*4 + 9fffch], eax
inc esi
loop column_loop
row bookkeeping
dec height
jnz row_loop
Because ESI starts at 1, the first destination is:
1 * 4 + 0x9fffc = 0xa0000
The column count 0x50 dwords is exactly one 320-byte scanline. This is a
visible-frame copy/expand stage: convert the software effect representation into
linear mode-13h bytes in VGA memory.
Routine 0x110cc4 is a related copy path, but instead of unpacking two nibbles
through the same table sequence, it shifts source dwords left by four and writes
them to the same 0xa0000 destination pattern. This gives another phase or
variant of the packed-buffer presentation.
Full-Screen Copy And Picture Parts
Several routines are plain, deliberate screen movers:
; around 0x11131b
mov esi, 114a34h
mov edi, 0a0000h
mov ecx, 3e80h
rep movsd
0x3e80 dwords is 0xFA00 bytes, exactly 64000 bytes. That is a full
320x200x8bpp screen copy.
Another path around 0x111351 copies rectangular regions with row skips:
copy some dwords
skip by 0xa0-byte style row increments
repeat for about 60 rows
This is the shape of a picture reveal, windowed blit, or scroll panel. It does not use ports; it relies on the custom VGA mode preserving a simple linear framebuffer.
The path around 0x1111d6 combines a screen copy with the CRTC scanline-height
change:
mov ax, 4309h
out 03d4h, ax
copy rows from 134284h+0140h to 134284h
set height/count state to 64h
That looks like a vertical scroll-buffer effect under altered hardware row height. It moves memory upward while the CRTC stretches rows differently, so the same software data can occupy a different apparent vertical space.
Frame Tick Update
Routine 0x110ec8 is the frame-state updater. It waits until a tick value near
0x117d4 changes, then updates many effect variables:
- Phase accumulators increase by small constants.
- Some values clamp at limits.
- Some values bounce by flipping an increment or branch direction.
0x12c134is incremented as a frame or table phase counter.
This routine is the bridge between music/timer synchronization and the visual loops. The renderers are tight because the per-frame state is prepared here.
Sin/Cos Vector Setup
Routines 0x110ffc and 0x1110e0 build vectors from the sine table and then
call renderers. The recurring pattern is:
phase_index = phase & 1ffch
sin_value = [112914h + phase_index]
cos_value = [112914h + ((phase_index + 0800h) & mask)]
The +0x800 offset is a quarter-wave offset in this table's indexing scheme,
so one table gives both sine and cosine. The values then become gradients or
increments for the mapper loops.
This is one of the reasons Xtal feels more polished than a simple tunnel demo. The mapper is not only scrolling a texture. It is being fed by moving vector terms, palette ramps, and offset tables that are all phase-aligned.
What The Major Parts Do
The exact source labels are gone, but the binary structure supports this functional map:
| Part | What it does | Evidence |
|---|---|---|
| Loader/runtime | Enters 32-bit PMODE/W, initializes Watcom runtime, opens XM and resources | PMODE/W marker, Watcom runtime strings, xtal.xm, MIDAS strings |
| VGA setup | Creates direct 320x200x256 framebuffer with 50 Hz timing | int 10h mode 13h followed by CRTC/SEQ/GC register block at 0x38a31a |
| Sync layer | Calibrates PIT against retrace and installs IRQ0 handler | Port 0x3da polling, PIT 0x43/0x40, vector 08h install |
| Music scheduler | Uses MIDAS/XM state as phase source for effect boundaries | 0x1100a4 pulls music-position fields and sections compare phase words |
| Palette system | Uploads whole DACs, builds ramps, performs hard white/black transitions | 0x3c8/0x3c9 loops and ramp builders |
| Picture blits | Expands RLE resources and copies full/rectangular screens | 0x1128ef, rep movsd 64000-byte copy |
| Crystal/tunnel mapper | Uses sine vectors and fixed-point texture addressing | 0x110cf3, texture at 0x143e57, two-pixel span loop |
| Wavy distortion | Builds wave offset tables and samples through shade maps | 0x110b6b, 0x110bf1, texture at 0x17550d, shade table at 0x12c144 |
| Packed display stage | Converts nibble/intermediate buffers to VGA bytes | 0x110c3b, 0x110c7d, writes to 0xa0000 |
| Scanline trick part | Changes vertical row height with CRTC register 0x09 |
writes 0x4309 and later 0x4109 to 0x3d4 |
The strongest interpretation is that Xtal is built around a normal linear mode-13h programming model, but with the hardware timing and display cadence changed underneath it. That gives the authors the speed of ordinary byte writes while still getting a distinctive 50 Hz display behavior and scanline-height effects.
Why The Inner Loops Are Good
The impressive part is not one huge algorithm. It is a stack of small, precise choices:
- The VGA mode keeps framebuffer writes linear.
- The 50 Hz timing is handled once in CRTC setup, not simulated in software.
- IRQ0 and retrace are tied together so frame work happens at stable moments.
- Music position is converted to compact phase values that all parts can test.
- Sine, shade, palette, and offset tables are built once per frame or once at startup.
- Texture loops use fixed-point high-byte extraction and packed 16-bit stores.
- Visible conversion to
0xa0000is a full scanline-friendly dword loop.
The renderers therefore spend their hot cycles on byte loads, table lookups, adds, shifts, and aligned stores. Expensive work is moved out of the pixel loop. That is the central engineering pattern in Xtal.
Runtime-To-Code Concordance
The six direct excerpts now map the complete visible route to the static analysis:
- The Complex logo and Xtal title map to PMODE/W/MIDAS setup, the custom 50 Hz
VGA mode at
0x38a31a, whole-DAC upload paths at0x1100f8/0x111340, RLE/resource expansion at0x1128ef, and the full-screen/rectangular copy paths. - The face, eye, starburst, and late-tunnel excerpts map to the
radial/crystal/tunnel renderer family: runtime sine generation at
0x11288f, shade/ramp construction at0x1128ab, per-frame offset tables at0x110b6b, distorted texture sampling at0x110bf1, and the symmetric/fixed-point mapper at0x110cf3. - The visible breathing center, rotating lattice, face/eye overlays, and palette changes match phase accumulators feeding table-driven texture and shade lookups rather than a sequence of prerecorded frames.
- The dense green/white finale maps to the packed-buffer-to-VGA path at
0x110c3b/0x110c7dand the direct0xa0000full-screen copy stages. - The stable
complex / xtal / jmagic - jugi - rewardcard followed by the natural DOS return proves that the MIDASNo Soundroute completes the authored scheduler and cleanup path.
Open Ends
I did not fully decompile the Watcom runtime, MIDAS internals, or every effect
transition. The low-level claims above are based on static disassembly of the
flattened PMODE/W image, constants and ports that are unambiguous in the
binary, and a synchronized complete direct capture. The remaining boundary is
exact music-position tracing: the article does not claim the precise
0x117dc..0x117e4 phase word for every captured frame.
The pieces that would be worth another pass are:
- A dynamic trace matching direct-capture frames to the music-position phase constants.
- Symbolic renaming of all effect-state variables around
0x1102a8..0x117e4. - A standalone recreation of the
0x38a31aVGA mode to verify exact horizontal and vertical frequencies on an emulator model. - A small reimplementation of
0x110cf3and0x110bf1against carved texture data, which would confirm the visual identity of each mapper path.