Skip to main content

Case study: BUG compiler

BUG is a small experimental language designed to demonstrate ethdebug/format integration. This case study explains how BUG implements debug information generation, providing a reference for other compiler authors.

Try it yourself

See the BUG playground to experiment with BUG and see its debug output.

BUG language overview​

BUG is a minimal smart contract language with:

  • Storage declarations — Named variables with explicit slot positions
  • Elementary types — uint256, bool, address, etc.
  • Composite types — Arrays, mappings, structs
  • Control flow — if, while, expressions
  • Two code sections — create (constructor) and code (runtime)

A simple BUG contract:

name Counter;

storage {
[0] count: uint256;
[1] threshold: uint256;
}

create {
count = 0;
threshold = 100;
}

code {
count = count + 1;
if (count >= threshold) {
count = 0;
}
}

Casts​

x as T converts x to type T as Solidity does. A cast to a narrower integer type or to address keeps the low bits; a cast to a signed type then sign-extends from the new width. A cast to a narrower bytesN keeps the leading bytes, and a cast to a wider one pads with zero bytes at the end. A cast from dynamic bytes to bytesN uses its first 32 bytes (zero past its length) as a bytes32. As in Solidity, dynamic bytes casts to an integer or address only through bytesN.

300 as uint8;                    // 44
200 as int8; // -56
(-1 as int256) as uint8; // 255
0x11223344 as bytes8; // 0x1122334400000000
0x1122334455667788 as bytes4; // 0x11223344
msg.data[0:4] as bytes4; // the function selector
msg.data[4:36] as bytes32 as address; // an ABI-encoded address

Signed operators​

Comparisons, /, and % on signed integers are signed, as in Solidity. Division truncates toward zero, and a remainder has the sign of the dividend. Arithmetic does not check for overflow.

Both operands of an arithmetic or comparison operator must be signed, or both unsigned; cast one to mix them. An integer literal takes the type of the other operand, if that type can hold it, so x < 0 and -1 < x work for a signed x, and x < 128 is an error for an int8.

(-7 as int8) < (2 as int8);      // true
(-7 as int8) / (2 as int8); // -3
(-7 as int8) % (2 as int8); // -1
(7 as int8) % (-2 as int8); // 1

Arithmetic​

Arithmetic wraps, as in a Solidity unchecked { } block; BUG has no overflow checks. A result narrower than 256 bits wraps to the width of its type, as a cast to that type does: an unsigned result keeps its low bits, and a signed result then sign-extends. An operation on two integer types has the wider type, and a literal takes the other operand's type, so x + 1 with x: uint8 is a uint8 and wraps at 8 bits.

(255 as uint8) + (1 as uint8);   // 0
(0 as uint8) - (1 as uint8); // 255
(16 as uint8) * (16 as uint8); // 0
(127 as int8) + (1 as int8); // -128
(-128 as int8) - (1 as int8); // 127
(-128 as int8) / (-1 as int8); // -128
-(-128 as int8); // -128

Bounds​

An index out of bounds reverts, as in Solidity. a[i] on an array or on dynamic bytes in memory, or on a fixed-size array in storage, reverts unless i is less than the length; this holds for reads and writes, and at each level of a nested array. A slice b[s:e] reverts unless s <= e <= b.length. An index or slice bound must be an unsigned integer; cast a signed one, as in a[i as uint256]. The revert data is Solidity's Panic(0x32), and the REVERT carries a revert context with panic: 50 and a pointer to that data. A dynamic array in storage is not checked: BUG has no push, so a write past the end is how such an array grows.

let a: array<uint256> = [1, 2, 3];
a[2] = 30; // ok
a[3] = 40; // reverts with Panic(0x32)
msg.data[0:4]; // reverts if the calldata is shorter

Compilation pipeline​

BUG uses a multi-stage compilation pipeline:

Source → AST → IR → EVM Bytecode
↓ ↓ ↓
Types Debug Program
  1. Parsing — Source text to AST with source locations
  2. Type checking — Validates types and collects type information
  3. IR generation — Converts AST to intermediate representation
  4. EVM code generation — Produces final bytecode with debug annotations

Debug information strategy​

BUG generates ethdebug/format output alongside bytecode by:

  1. Tracking source locations through all compilation phases
  2. Preserving type information from the type checker
  3. Computing storage layouts during IR generation
  4. Emitting program annotations during code generation

Key design decisions​

Type preservation: BUG's IR types carry an "origin" field linking back to the source-level type. This allows generating rich ethdebug/format type information even after type erasure in the IR.

Storage analysis: During IR generation, BUG analyzes storage access patterns to determine variable locations. This analysis traces through compute_slot instructions to reconstruct the storage layout.

Program builder: A dedicated ProgramBuilder class accumulates instructions with their contexts during code generation, then serializes the complete program annotation.

Type generation​

BUG converts its type system to ethdebug/format types:

packages/bugc/src/irgen/debug/types.ts
export function convertBugType(bugType: BugType): Format.Type | undefined {
// Elementary types
if (BugType.isElementary(bugType)) {
return convertElementaryType(bugType);
}

// Array types
if (BugType.isArray(bugType)) {
const elementType = convertBugType(bugType.element);
return {
kind: "array",
contains: { type: elementType },
...(bugType.size !== undefined && { count: bugType.size }),
};
}

// Mapping and struct types follow similar patterns...
}

The type conversion handles:

  • Elementary types — Direct mapping (uint256 → {kind: "uint", bits: 256})
  • Arrays — Recursive conversion with optional count
  • Mappings — Key/value type conversion
  • Structs — Field-by-field conversion with names

Pointer generation​

BUG generates pointers that describe how to locate variables at runtime.

Storage variables​

For simple storage variables, BUG generates direct slot pointers:

{
"location": "storage",
"slot": 0
}

Composite types​

For structs, BUG generates group pointers with field offsets:

{
"group": [
{ "name": "field1", "location": "storage", "slot": 0 },
{ "name": "field2", "location": "storage", "slot": 0, "offset": 16 }
]
}

Dynamic arrays​

For dynamic arrays, BUG generates pointers with keccak256 expressions:

{
"group": [
{ "name": "array-length", "location": "storage", "slot": 0 },
{
"list": {
"count": { "$read": "array-length" },
"each": "i",
"is": {
"name": "element",
"location": "storage",
"slot": { "$sum": [{ "$keccak256": [{ "$wordsized": 0 }] }, "i"] }
}
}
}
]
}

Storage analysis​

BUG includes a storage analysis pass that traces compute_slot instructions to reconstruct dynamic storage locations. This handles patterns like:

compute_slot(mapping, baseSlot, key) → keccak256(key, slot)
compute_slot(array, baseSlot) → keccak256(slot)
compute_slot(field, baseSlot, offset) → slot + offset

Program annotation​

BUG emits program annotations that map bytecode to source context.

Instruction context​

Each bytecode instruction includes context with:

  • Source range — Where in source this instruction originates
  • Variables — Variables in scope at this point
  • Remarks — Human-readable annotations
{
"offset": "0x1a",
"operation": { "mnemonic": "SLOAD" },
"context": {
"gather": [
{
"code": {
"source": { "id": "main" },
"range": { "offset": 120, "length": 5 }
}
},
{
"variables": [
{
"identifier": "count",
"type": { "kind": "uint", "bits": 256 },
"pointer": { "location": "storage", "slot": 0 }
}
]
}
]
}
}

Program builder​

The ProgramBuilder class manages instruction accumulation:

Simplified pattern
class ProgramBuilder {
private instructions: Program.Instruction[] = [];

addInstruction(
offset: number,
operation: Program.Instruction.Operation,
context?: Program.Context,
) {
this.instructions.push({ offset, operation, context });
}

build(): Program {
return {
instructions: this.instructions,
};
}
}

Variable scoping​

BUG includes storage variables in the variables context for every instruction that accesses them. This allows debuggers to inspect storage values at any execution point.

packages/bugc/src/irgen/debug/variables.ts
export function collectVariablesWithLocations(
state: State,
sourceId: string,
): VariableInfo[] {
const variables: VariableInfo[] = [];

// Storage variables have fixed slots
for (const storageDecl of state.module.storageDeclarations) {
const bugType = state.types.get(storageDecl.id);
const pointer = generateStoragePointer(storageDecl.slot, bugType);

variables.push({
identifier: storageDecl.name,
type: convertBugType(bugType),
pointer,
declaration: storageDecl.loc
? {
source: { id: sourceId },
range: storageDecl.loc,
}
: undefined,
});
}

return variables;
}

BUG also lists the local variables (parameters and lets) in scope at each instruction, each with a pointer to its value where the value is located: a memory word in the function's frame, or a stack slot. A local hides a storage variable of the same name while it is in scope. Under optimization, a local whose value the optimizer folded to a constant or removed is listed without a pointer. See packages/bugc/src/evmgen/debug/local-variables.ts.

Testing strategy​

BUG tests debug output in several ways:

  1. Unit tests — Test individual conversion functions
  2. Integration tests — Compile programs and verify output structure
  3. Playground — Visual verification of output (see BUG playground)

Example test pattern:

test("generates correct storage pointer", () => {
const source = `
name Test;
storage { [0] value: uint256; }
code { value = 42; }
`;

const result = compile({ to: "bytecode", source });
const program = result.value.program;

// Find the SLOAD instruction
const sload = program.instructions.find(
(i) => i.operation?.mnemonic === "SLOAD",
);

expect(sload.context).toMatchObject({
variables: [
{
identifier: "value",
pointer: { location: "storage", slot: 0 },
},
],
});
});

Lessons learned​

Start simple​

BUG started with just storage variables and elementary types. Complex features (arrays, mappings, structs) were added incrementally after the basic infrastructure was working.

Preserve information early​

Type information is easier to preserve than reconstruct. BUG's IR types carry their source-level origin, which simplifies later conversion.

Test visually​

The BUG playground proved invaluable for debugging the debug output. Being able to see the generated annotations alongside bytecode helped catch issues that unit tests missed.

Handle edge cases gracefully​

When storage analysis can't determine a location (e.g., computed slot from a non-constant), BUG still generates useful partial information rather than failing entirely.

Future work​

Areas for improvement in BUG's debug support:

  • Richer contexts — Add frame information for function calls
  • Source maps — More granular source location tracking

Resources​