Analysis model: gpt-5.5 xhigh

Xtal by Complex Media Labs - Technical Dissection

Scope

This is a binary and complete direct-runtime dissection of xtal - the demonstration by Finnish group Complex Media Labs, released in 1995. The public production metadata identifies it as a PC demo, and the release text inside the archive credits several important runtime components:

The caveat is important: this is not source code. The executable is a packed PMODE/W protected-mode binary. I flattened the PMODE/W image, inspected the recovered flat memory map, and ran the authenticated executable through its complete native no-sound route. A few PMODE/W relocation warnings were emitted while flattening, so control-flow and local loops are reliable enough for analysis, but any claim about high-level symbol names should be read as an inference from code shape, constants, data flow, and direct runtime concordance.

Examined Files

The archive contains:

XTAL.EXE    PMODE/W-packed DOS executable
XTAL.XM     FastTracker II module, title "gates of ixtlan"
XTAL.TXT    release notes
file_id.diz BBS description

Hashes from the examined files:

349571136dd60871d9c4af32b278cedd2a46bb49879f3d8566f59cc6463d49f6  clx_xtal.zip
a31e3d1ec52d018158935c2665f86701b36da54ac322f629a61d09f06f8dc87d  XTAL.EXE
b9163e835e645edb33b6bc73b57c398ac8e06c137c798ef07e7212fc863dd9ad  XTAL.XM

XTAL.TXT gives the most useful non-code hints:

- doesnt work with win95 :)
- stereo mode is forced to mono, except on gus...
- music player is midas v0.50 pre-release #1 by sahara surfers
- dos extender is pmode/w v1.21 by charles scheffold and thomas pytel
- sine tables are runtime calculated using karl's sine generator
- 320x200x256 50hz mode is reverse-engineered by liket

file_id.diz describes the release as the final 4 MB memory version.

Complete Direct Silent Runtime

The earlier runtime pass used DOSBox-X's internal H.264 recorder, selected Gravis UltraSound, stopped on a schedule after 234.6 seconds, and then killed DOSBox. It showed useful imagery but did not prove the endpoint.

The new run uses the authenticated archive and MIDAS's own No Sound device. No guest sound card is selected, SDL uses a dummy host driver, and the external FFV1 master contains video only. The visual scheduler advances normally and the executable returns naturally after its final credit card.

Emulator:          DOSBox-X 2026.01.02
Machine:           S3 VGA, 16 MB RAM
CPU:               normal 486 core, fixed 50,000 cycles
Memory:            XMS on; EMS and UMB off
Show path:         original XTAL.EXE, MIDAS "No Sound"
Capture:           external X11 grab, 25 fps FFV1, no audio stream
Working master:    226.2 seconds including setup
Endpoint:          natural return to DOS; DOSBox exit status 0
Published media:   six continuous GIF excerpts; no audio or H.264/MP4

Approximate direct-capture phases:

Capture time Directly observed state
00:00–00:11 MIDAS setup and the selected No Sound path.
00:12–00:32 Handwritten Complex logo, production card, and spaced xtal title.
00:33–00:52 Magenta and green radial bursts establish the shared tunnel field.
00:53–01:13 Painted face and luminous wire/mesh layers blend into the tunnel.
01:14–01:43 Face, eye and concentric material repeatedly enter the same radial transform.
01:44–02:43 Bright starbursts, red rings and dense high-contrast tunnel phases.
02:44–03:26 Green, grey and white contour tunnels become increasingly warped and dense.
03:27–03:46 Final tunnel dissolves into the complex / xtal / jmagic - jugi - reward card, then the executable returns to DOS.

These are continuous excerpts from that single direct silent run:

Xtal handwritten Complex opening logo

Xtal radial field blending into the painted face and mesh sequence

Xtal eye and concentric radial transformation

Xtal bright starburst and tunnel motion

Xtal dense green-white late tunnel

Xtal final tunnel transition toward the closing credit card

Executable Shape

The MZ header describes a small real-mode loader followed by a PMODE/W payload:

XTAL.EXE size:        1141351 bytes
MZ image size:        9808 bytes
header paragraphs:   4
relocations:         0
initial CS:IP:       0000:0059
initial SS:SP:       0344:0100

The string PMODE/W v1.21 occurs in the real-mode stub, and the PMODE/W packed payload marker PMW1 appears at file offset 0x2650, immediately after the MZ loader image.

Flattening the PMODE/W image produced:

Flat memory map: /tmp/xtal-analysis/unpacked/XTAL.EXE.FLAT
Flat size:       3782003 bytes
Entry point:     0x003909D8
Stack pointer:   0x00012A90

The flat image contains ordinary Watcom runtime text near the entry point:

WATCOM C/C++32 Run-Time system.

That is consistent with a 32-bit protected-mode C or C++ program using runtime support code, direct hardware routines, and hand-written assembly inner loops.

High-Level Program Flow

At 0x3909d8, execution jumps over the Watcom copyright string into startup code. The early path is mostly runtime and DOS-extender setup: stack, argument environment, selectors, file and memory setup, then device and music setup.

The demo-specific setup path calls these important routines:

0x38a31a  custom VGA mode setup
0x390646  retrace/PIT calibration
0x390739  IRQ0 vector installation path
0x11288f  sine-table recurrence builder
0x1128ab  shade/ramp table builder

The archive also stores the music as a separate XTAL.XM file. The executable contains strings such as:

xtal.xm
MIDAS Error
Using %s
Extended Module
ULTRASND
BLASTER

So the binary loads the external XM module through MIDAS, probes common DOS audio configuration paths, and uses the music player's reported position as part of its effect scheduler.

The central scheduler repeatedly does three things:

  1. Run the current effect update or renderer.
  2. Poll the keyboard with BIOS int 16h, ah=1 and branch to exit if needed.
  3. Call the timing/music-position updater at 0x1100a4 and compare the resulting clock against phase thresholds.

The common phase value used by many sections is built from fields near 0x117dc..0x117e4. A typical compare loads:

movzx eax, byte [0x117dc]
shl   eax, 8
or    al, [0x117e4]
cmp   ax, phase_limit

That makes the demo timeline music-driven rather than just frame-count driven.

The 50 Hz VGA Mode

The most explicit hardware trick is the custom 320x200x256 50 Hz VGA mode at 0x38a31a. It starts from BIOS mode 13h and then rewrites VGA timing registers:

mov ax, 0013h
int 10h

Then it unlocks protected CRTC registers:

mov dx, 03d4h
mov al, 11h
out dx, al
inc dl
in  al, dx
and al, 7fh
out dx, al

CRTC register 0x11 is the vertical retrace end register. Bit 7 is the CRTC write-protect bit. Clearing it allows writes to the vertical timing registers.

Next the code writes the VGA miscellaneous output register:

mov dl, 0c2h
mov al, 0e3h
out dx, al

Then it returns to the color CRTC index port and writes a full timing block. The code uses 16-bit indexed VGA writes: AL is the register index, AH is the value, and out dx, ax writes index to 0x3d4 and value to 0x3d5.

CRTC 00 = 60
CRTC 01 = 4f
CRTC 02 = 50
CRTC 03 = 82
CRTC 04 = 54
CRTC 05 = 80
CRTC 06 = 6f
CRTC 07 = 3e
CRTC 08 = 00
CRTC 09 = 41
CRTC 10 = f6
CRTC 11 = 88
CRTC 12 = 8f
CRTC 13 = 28
CRTC 14 = 40
CRTC 15 = 90
CRTC 16 = 6c
CRTC 17 = a3

The key consequences:

The horizontal setup is also not a vanilla BIOS table:

After CRTC setup, the code programs the sequencer:

SEQ 1 = 01
SEQ 3 = 00
SEQ 4 = 0e

SEQ 4 = 0x0e enables the mode-13h style memory behavior: extended memory, odd/even disabled, and chain-4 addressing.

Then it programs the graphics controller:

GC 5 = 40
GC 6 = 05

GC 5 = 0x40 selects the 256-color shift behavior, and GC 6 = 0x05 selects graphics mode with the A0000 aperture. The result is deliberately practical: the video memory is still easy to address as a linear 64000-byte surface, but the monitor timing has been bent to 50 Hz.

That is why the release note says the mode was reverse-engineered and does not work under Windows 95. This code assumes direct VGA I/O, a real retrace status bit, and hardware that accepts non-BIOS timing values.

Scanline Height Trick

The demo changes CRTC register 0x09, the maximum scanline register, in later parts:

; around 0x1111d6
mov dx, 03d4h
mov ax, 4309h
out dx, ax

; around 0x112850
mov dx, 03d4h
mov ax, 4109h
out dx, ax

The normal custom mode sets CRTC 09 = 0x41; this later code temporarily uses 0x43. The low bits of this register control scanlines per character row. In a 256-color chained mode this changes the vertical stretching and row cadence without changing the byte layout of the software buffers. That is a classic VGA-screen trick: one effect can be made taller, chunkier, or scroller-like by changing hardware row repetition instead of resampling the whole image.

Retrace And Timer Synchronization

The binary installs its own IRQ0 handler. The vector-management path around 0x390739 uses DOS interrupt vector calls:

mov ax, 3508h
int 21h        ; get old IRQ0 vector

mov edx, 390392h
mov ax, 2508h
int 21h        ; install new IRQ0 handler

The handler at 0x390392 saves all general registers and segment registers, sets its data segments, and then branches on a state byte around 0xfc2e.

The interesting mode waits on the VGA input-status register:

mov dx, 03dah

wait_until_not_in_retrace:
    in   al, dx
    test al, 08h
    jne  wait_until_not_in_retrace

; optional callback through [0xfc16]
; fixed-point counters updated here

wait_until_in_retrace:
    in   al, dx
    test al, 08h
    je   wait_until_in_retrace

; optional callback through [0xfc1a]
call 3902ddh

mov al, 20h
out 20h, al     ; PIC end-of-interrupt

; optional callback through [0xfc1e]

So IRQ0 is not just a plain timer tick. It is interlocked with the VGA vertical retrace bit. That lets the demo run frame callbacks at stable display moments, update counters, and avoid tearing-sensitive updates happening in the middle of visible scanout.

The calibration routine at 0x390646 measures the retrace period using the PIT:

; wait for a retrace transition through port 03dah
mov al, 36h
out 43h, al      ; PIT channel 0, mode 3, lobyte/hibyte
xor al, al
out 40h, al
out 40h, al      ; reload counter with 0

; wait for another retrace transition
mov al, 00h
out 43h, al      ; latch counter 0
in  al, 40h
mov ah, al
in  al, 40h
xchg al, ah
neg ax           ; elapsed count

It repeats the measurement and accepts it only when two readings differ by at most two PIT ticks. That is a good tell that the authors wanted a stable frame duration, not a rough delay.

Palette And DAC Paths

Xtal spends a lot of work on palette control. The direct DAC upload routine at 0x1100f8 writes a whole 256-color palette:

mov edx, 03c8h
xor al, al
out dx, al          ; DAC write index = 0

mov ebx, palette
lea ecx, [ebx+0300h]
mov esi, 03c9h

upload:
    mov al, [ebx+0]
    mov edx, esi
    out dx, al
    mov al, [ebx+1]
    out dx, al
    mov al, [ebx+2]
    add ebx, 3
    out dx, al
    cmp ebx, ecx
    jne upload

That is exactly 0x300 bytes: 256 RGB triples. It sets the start index to zero with port 0x3c8, then writes RGB data to port 0x3c9.

There is a second compact DAC upload at 0x111340:

mov dx, 03c8h
xor al, al
out dx, al
inc dl             ; 03c9h
mov ecx, 0300h
rep outsb

Two special routines force the whole DAC to white or black:

; whiteout around 0x112798
mov dx, 03c8h
xor al, al
out dx, al
inc dl
not al             ; ffh
mov ecx, 0300h
rep outsb-like loop

; blackout around 0x11287d
mov dx, 03c8h
xor al, al
out dx, al
inc dl
mov ecx, 0300h
rep zero output

Classic VGA DACs store 6-bit color values, so values above 0x3f are effectively saturated or masked by the hardware. The intent is still clear: instant full-screen white and black transitions without touching the framebuffer.

Resource Expansion

The image/resource unpacker at 0x1128ef is a tiny RLE decoder. It receives a destination in EDI and an output byte count in ECX, then computes the end pointer:

lea edx, [edi+ecx]

The stream is controlled by one byte at a time:

next_packet:
    mov  cl, [esi]
    inc  esi
    shr  cl, 1
    jc   run_packet

literal_packet:
    mov  al, cl
    stosb
    cmp  edi, edx
    jb   next_packet
    ret

run_packet:
    mov  al, [esi]
    inc  esi
    rep  stosb
    cmp  edi, edx
    jb   next_packet
    ret

Odd control bytes mean "repeat the following byte control >> 1 times". Even control bytes mean "emit control >> 1 as a literal byte". This is small and fast, and it fits palette-index images well because the literal path does not need to carry a second source byte.

The unpacker is used for both full-screen and off-screen buffers. One path expands directly to 0xa0000, and others expand into software buffers that are later transformed or copied.

Runtime Sine And Shade Tables

The release note about Karl's sine generator matches the code. Routine 0x11288f builds a recurrence table at runtime:

mov edi, 11291ch
mov ecx, 07feh

loop:
    mov ebx, [edi-4]
    mov eax, ebx
    imul ebx
    shrd eax, edx, 1dh
    sub eax, [edi-8]
    stosd
    loop loop

This is not calling sin() from a library. It is using previous table values to generate the next ones. The code then uses masked phase indexes such as phase & 0x1ffc, and a +0x800 phase offset for the perpendicular component. That gives sine/cosine pairs from one table.

Routine 0x1128ab builds a shade or distance ramp table near 0x12c144. Its outer loops walk signed byte-like coordinates, and its inner part computes and clamps a brightness:

; simplified shape
for bh from 7fh down through signed range:
  for bl from 7fh down through signed range:
      value = ecx + edx + 3fh
      if high_byte(value) > 78h:
          value = 7800h
      store high_byte(value)

The exact table consumers show that this is used as a mapping table during texture distortion. A texture sample is not written directly; it is combined with a computed index and then looked up through this precomputed ramp. That keeps the inner loop byte-sized and avoids multiplies or square roots per pixel.

Palette Ramp Builder

The helper at 0x1107ef writes a 64-entry linear ramp. It takes a starting fixed-point value and an increment, clamps the increment to 0..0x100, then stores the high byte for 64 entries:

cmp edx, 100h
jbe increment_ok
mov edx, 100h

increment_ok:
test eax, eax
js   negative_start

mov ecx, 40h
positive_loop:
    stosb_as_ah
    add eax, edx
    loop positive_loop

The negative-start path emits zero until the accumulator becomes positive, then continues with the same high-byte ramp. This is exactly the kind of primitive you want for palette fades and color-table morphs: it turns a phase and slope into a compact 64-byte channel curve.

Higher-level routines combine these curves:

0x110828  builds a 16x16-ish table from channel ramps and a bias
0x1109e8  combines two 0x300-byte palette/table blocks into 0x114d0
0x110a20  drives palette/table morphing from phase variables

The important point is that palette motion is not just "fade current DAC toward black". The demo builds intermediate tables and uses them in render loops. That is why the color changes can be tied to distortion and shading, not only to global brightness.

Distortion Offset Table

Routine 0x110b6b builds a 256-entry word table at 0x110329. It reads the runtime sine table, uses a small function table around 0x110bb0, and writes per-line or per-column offsets:

phase = [0x1102a8] + frame_dependent_terms
for i in 0..255:
    sample = sine_table[(phase + i * step) & mask]
    offset_table[i] = transformed(sample)

Several variables around 0x1102a8, 0x1102b0, and related addresses feed this builder. These are not static lookup tables baked into the executable; the waves are recomputed as phase accumulators move.

This table is then consumed by the texture distortion path. The pattern is classic for mid-1990s VGA demo effects:

  1. Update wave phases once per frame.
  2. Build a small offset table.
  3. Use that table inside a byte-per-pixel loop to avoid expensive math.

Distorted Texture Sampler

Routine 0x110bf1 is one of the important visible-effect loops. It reads an index table, adds a wave displacement, samples a texture, maps it through a shade table, and writes an off-screen byte buffer.

The setup is roughly:

edi = 124434h          ; destination buffer
ecx = height * 0a0h - 1

The hot path is:

loop:
    ; fetch base index from a word table
    movzx eax, word [index_table + ...]

    ; add dynamic wave displacement
    add ax, [110329h + ...]

    ; sample texture
    movzx eax, byte [eax + 17550dh]

    ; shade/remap sample
    mov al, [eax + 12c144h]

    ; write software framebuffer byte
    stosb

    dec ecx
    jnz loop

This loop is important because it separates the expensive-looking visual from the real work:

That is why the effect can run on mid-1990s PC hardware. The inner loop does no division, no trigonometry, and no per-pixel port I/O.

Symmetric Texture Mapper

Routine 0x110cf3 is a denser mapper. It renders two pixels per iteration from a source texture around 0x143e57, using fixed-point u/v accumulators and then draws a mirrored half.

The hot two-pixel part looks like this:

mov eax, 004f0000h

span_loop:
    mov ebx, ecx
    shr ebx, 24
    add ecx, esi
    shld ebx, edx, 8
    add edi, 2

    mov al, [ebx + 143e57h]
    add edx, ebp

    mov ebx, ecx
    shr ebx, 24
    add ecx, esi
    shld ebx, edx, 8

    mov ah, [ebx + 143e57h]
    add edx, ebp

    mov [edi-2], ax
    sub eax, 00010000h
    jae span_loop

What this does:

After the forward span, the routine reverses direction for the mirrored half by subtracting the accumulators and writing backward. That matters: the effect is not rendering every pixel independently. It is using symmetry to halve setup work and keep the span loop tight.

The texture address construction is also a neat trick. Instead of computing y * width + x, the texture is arranged so that selected high bits from two fixed-point accumulators can be fused into a usable address. That is a very demo-scene way to trade memory layout for speed.

Nibble Compression And VRAM Expansion

The software effect buffers are not always stored in final VGA form. Routine 0x110c3b processes 0x1f40 dwords with a mask:

mov edx, 0f0f0f0fh
mov ecx, 1f40h

loop:
    mov eax, [esi]
    and eax, edx
    ; optional shift by 4 depending on [0x1102f8]
    mov [esi], eax
    add esi, 4
    loop loop

0x1f40 dwords is 32000 bytes. The 0x0f0f0f0f mask keeps low nibbles from each byte. The optional shift selects the other nibble phase. This is a cheap way to prepare a half-resolution, 4-bit, or interleaved-looking buffer before the final VGA write.

Routine 0x110c7d expands intermediate data into actual VGA memory near 0xa0000. The loop uses a base table around 0x114a30, reads packed nibbles from 0x124434, and writes dwords to the framebuffer:

esi = 1
height = [0x1107eb]

row_loop:
    ecx = 50h              ; 80 dwords = 320 bytes

    column_loop:
        ; unpack or map two nibbles through a table
        ; build one 32-bit group of VGA palette indexes
        mov [esi*4 + 9fffch], eax
        inc esi
        loop column_loop

    row bookkeeping
    dec height
    jnz row_loop

Because ESI starts at 1, the first destination is:

1 * 4 + 0x9fffc = 0xa0000

The column count 0x50 dwords is exactly one 320-byte scanline. This is a visible-frame copy/expand stage: convert the software effect representation into linear mode-13h bytes in VGA memory.

Routine 0x110cc4 is a related copy path, but instead of unpacking two nibbles through the same table sequence, it shifts source dwords left by four and writes them to the same 0xa0000 destination pattern. This gives another phase or variant of the packed-buffer presentation.

Full-Screen Copy And Picture Parts

Several routines are plain, deliberate screen movers:

; around 0x11131b
mov esi, 114a34h
mov edi, 0a0000h
mov ecx, 3e80h
rep movsd

0x3e80 dwords is 0xFA00 bytes, exactly 64000 bytes. That is a full 320x200x8bpp screen copy.

Another path around 0x111351 copies rectangular regions with row skips:

copy some dwords
skip by 0xa0-byte style row increments
repeat for about 60 rows

This is the shape of a picture reveal, windowed blit, or scroll panel. It does not use ports; it relies on the custom VGA mode preserving a simple linear framebuffer.

The path around 0x1111d6 combines a screen copy with the CRTC scanline-height change:

mov ax, 4309h
out 03d4h, ax

copy rows from 134284h+0140h to 134284h
set height/count state to 64h

That looks like a vertical scroll-buffer effect under altered hardware row height. It moves memory upward while the CRTC stretches rows differently, so the same software data can occupy a different apparent vertical space.

Frame Tick Update

Routine 0x110ec8 is the frame-state updater. It waits until a tick value near 0x117d4 changes, then updates many effect variables:

This routine is the bridge between music/timer synchronization and the visual loops. The renderers are tight because the per-frame state is prepared here.

Sin/Cos Vector Setup

Routines 0x110ffc and 0x1110e0 build vectors from the sine table and then call renderers. The recurring pattern is:

phase_index = phase & 1ffch
sin_value   = [112914h + phase_index]
cos_value   = [112914h + ((phase_index + 0800h) & mask)]

The +0x800 offset is a quarter-wave offset in this table's indexing scheme, so one table gives both sine and cosine. The values then become gradients or increments for the mapper loops.

This is one of the reasons Xtal feels more polished than a simple tunnel demo. The mapper is not only scrolling a texture. It is being fed by moving vector terms, palette ramps, and offset tables that are all phase-aligned.

What The Major Parts Do

The exact source labels are gone, but the binary structure supports this functional map:

Part What it does Evidence
Loader/runtime Enters 32-bit PMODE/W, initializes Watcom runtime, opens XM and resources PMODE/W marker, Watcom runtime strings, xtal.xm, MIDAS strings
VGA setup Creates direct 320x200x256 framebuffer with 50 Hz timing int 10h mode 13h followed by CRTC/SEQ/GC register block at 0x38a31a
Sync layer Calibrates PIT against retrace and installs IRQ0 handler Port 0x3da polling, PIT 0x43/0x40, vector 08h install
Music scheduler Uses MIDAS/XM state as phase source for effect boundaries 0x1100a4 pulls music-position fields and sections compare phase words
Palette system Uploads whole DACs, builds ramps, performs hard white/black transitions 0x3c8/0x3c9 loops and ramp builders
Picture blits Expands RLE resources and copies full/rectangular screens 0x1128ef, rep movsd 64000-byte copy
Crystal/tunnel mapper Uses sine vectors and fixed-point texture addressing 0x110cf3, texture at 0x143e57, two-pixel span loop
Wavy distortion Builds wave offset tables and samples through shade maps 0x110b6b, 0x110bf1, texture at 0x17550d, shade table at 0x12c144
Packed display stage Converts nibble/intermediate buffers to VGA bytes 0x110c3b, 0x110c7d, writes to 0xa0000
Scanline trick part Changes vertical row height with CRTC register 0x09 writes 0x4309 and later 0x4109 to 0x3d4

The strongest interpretation is that Xtal is built around a normal linear mode-13h programming model, but with the hardware timing and display cadence changed underneath it. That gives the authors the speed of ordinary byte writes while still getting a distinctive 50 Hz display behavior and scanline-height effects.

Why The Inner Loops Are Good

The impressive part is not one huge algorithm. It is a stack of small, precise choices:

The renderers therefore spend their hot cycles on byte loads, table lookups, adds, shifts, and aligned stores. Expensive work is moved out of the pixel loop. That is the central engineering pattern in Xtal.

Runtime-To-Code Concordance

The six direct excerpts now map the complete visible route to the static analysis:

Open Ends

I did not fully decompile the Watcom runtime, MIDAS internals, or every effect transition. The low-level claims above are based on static disassembly of the flattened PMODE/W image, constants and ports that are unambiguous in the binary, and a synchronized complete direct capture. The remaining boundary is exact music-position tracing: the article does not claim the precise 0x117dc..0x117e4 phase word for every captured frame.

The pieces that would be worth another pass are: