Case study: BUG compiler
BUG is a small experimental language designed to demonstrate ethdebug/format integration. This case study explains how BUG implements debug information generation, providing a reference for other compiler authors.
See the BUG playground to experiment with BUG and see its debug output.
BUG language overview
BUG is a minimal smart contract language with:
- Storage declarations — Named variables with explicit slot positions
- Elementary types —
uint256,bool,address, etc. - Composite types — Arrays, mappings, structs
- Control flow —
if,while, expressions - Two code sections —
create(constructor) andcode(runtime)
A simple BUG contract:
name Counter;
storage {
[0] count: uint256;
[1] threshold: uint256;
}
create {
count = 0;
threshold = 100;
}
code {
count = count + 1;
if (count >= threshold) {
count = 0;
}
}
Casts
x as T converts x to type T as Solidity does. A cast to a narrower
integer type or to address keeps the low bits; a cast to a signed type
then sign-extends from the new width. A cast to a narrower bytesN keeps
the leading bytes, and a cast to a wider one pads with zero bytes at the
end. A cast from dynamic bytes to bytesN uses its first 32 bytes
(zero past its length) as a bytes32. As in Solidity, dynamic bytes
casts to an integer or address only through bytesN.
300 as uint8; // 44
200 as int8; // -56
(-1 as int256) as uint8; // 255
0x11223344 as bytes8; // 0x1122334400000000
0x1122334455667788 as bytes4; // 0x11223344
msg.data[0:4] as bytes4; // the function selector
msg.data[4:36] as bytes32 as address; // an ABI-encoded address
Signed operators
Comparisons, /, and % on signed integers are signed, as in Solidity.
Division truncates toward zero, and a remainder has the sign of the
dividend. Arithmetic does not check for overflow.
Both operands of an arithmetic or comparison operator must be signed, or
both unsigned; cast one to mix them. An integer literal takes the type of
the other operand, if that type can hold it, so x < 0 and -1 < x work
for a signed x, and x < 128 is an error for an int8.
(-7 as int8) < (2 as int8); // true
(-7 as int8) / (2 as int8); // -3
(-7 as int8) % (2 as int8); // -1
(7 as int8) % (-2 as int8); // 1
Arithmetic
Arithmetic wraps, as in a Solidity unchecked { } block; BUG has no
overflow checks. A result narrower than 256 bits wraps to the width of
its type, as a cast to that type does: an unsigned result keeps its low
bits, and a signed result then sign-extends. An operation on two integer
types has the wider type, and a literal takes the other operand's type,
so x + 1 with x: uint8 is a uint8 and wraps at 8 bits.
(255 as uint8) + (1 as uint8); // 0
(0 as uint8) - (1 as uint8); // 255
(16 as uint8) * (16 as uint8); // 0
(127 as int8) + (1 as int8); // -128
(-128 as int8) - (1 as int8); // 127
(-128 as int8) / (-1 as int8); // -128
-(-128 as int8); // -128
Bounds
An index out of bounds reverts, as in Solidity. a[i] on an array or on
dynamic bytes in memory, or on a fixed-size array in storage, reverts
unless i is less than the length; this holds for reads and writes, and
at each level of a nested array. A slice b[s:e] reverts unless
s <= e <= b.length. An index or slice bound must be an unsigned
integer; cast a signed one, as in a[i as uint256]. The revert data is Solidity's Panic(0x32), and
the REVERT carries a revert context with panic: 50 and a pointer to
that data. A dynamic array in storage is not checked: BUG has no push,
so a write past the end is how such an array grows.
let a: array<uint256> = [1, 2, 3];
a[2] = 30; // ok
a[3] = 40; // reverts with Panic(0x32)
msg.data[0:4]; // reverts if the calldata is shorter
Compilation pipeline
BUG uses a multi-stage compilation pipeline:
Source → AST → IR → EVM Bytecode
↓ ↓ ↓
Types Debug Program
- Parsing — Source text to AST with source locations
- Type checking — Validates types and collects type information
- IR generation — Converts AST to intermediate representation
- EVM code generation — Produces final bytecode with debug annotations
Debug information strategy
BUG generates ethdebug/format output alongside bytecode by:
- Tracking source locations through all compilation phases
- Preserving type information from the type checker
- Computing storage layouts during IR generation
- Emitting program annotations during code generation
Key design decisions
Type preservation: BUG's IR types carry an "origin" field linking back to the source-level type. This allows generating rich ethdebug/format type information even after type erasure in the IR.
Storage analysis: During IR generation, BUG analyzes storage access
patterns to determine variable locations. This analysis traces through
compute_slot instructions to reconstruct the storage layout.
Program builder: A dedicated ProgramBuilder class accumulates
instructions with their contexts during code generation, then serializes the
complete program annotation.
Type generation
BUG converts its type system to ethdebug/format types:
export function convertBugType(bugType: BugType): Format.Type | undefined {
// Elementary types
if (BugType.isElementary(bugType)) {
return convertElementaryType(bugType);
}
// Array types
if (BugType.isArray(bugType)) {
const elementType = convertBugType(bugType.element);
return {
kind: "array",
contains: { type: elementType },
...(bugType.size !== undefined && { count: bugType.size }),
};
}
// Mapping and struct types follow similar patterns...
}
The type conversion handles:
- Elementary types — Direct mapping (
uint256→{kind: "uint", bits: 256}) - Arrays — Recursive conversion with optional count
- Mappings — Key/value type conversion
- Structs — Field-by-field conversion with names
Pointer generation
BUG generates pointers that describe how to locate variables at runtime.
Storage variables
For simple storage variables, BUG generates direct slot pointers:
{
"location": "storage",
"slot": 0
}
Composite types
For structs, BUG generates group pointers with field offsets:
{
"group": [
{ "name": "field1", "location": "storage", "slot": 0 },
{ "name": "field2", "location": "storage", "slot": 0, "offset": 16 }
]
}
Dynamic arrays
For dynamic arrays, BUG generates pointers with keccak256 expressions:
{
"group": [
{ "name": "array-length", "location": "storage", "slot": 0 },
{
"list": {
"count": { "$read": "array-length" },
"each": "i",
"is": {
"name": "element",
"location": "storage",
"slot": { "$sum": [{ "$keccak256": [{ "$wordsized": 0 }] }, "i"] }
}
}
}
]
}
Storage analysis
BUG includes a storage analysis pass that traces compute_slot instructions
to reconstruct dynamic storage locations. This handles patterns like:
compute_slot(mapping, baseSlot, key) → keccak256(key, slot)
compute_slot(array, baseSlot) → keccak256(slot)
compute_slot(field, baseSlot, offset) → slot + offset
Program annotation
BUG emits program annotations that map bytecode to source context.
Instruction context
Each bytecode instruction includes context with:
- Source range — Where in source this instruction originates
- Variables — Variables in scope at this point
- Remarks — Human-readable annotations
{
"offset": "0x1a",
"operation": { "mnemonic": "SLOAD" },
"context": {
"gather": [
{
"code": {
"source": { "id": "main" },
"range": { "offset": 120, "length": 5 }
}
},
{
"variables": [
{
"identifier": "count",
"type": { "kind": "uint", "bits": 256 },
"pointer": { "location": "storage", "slot": 0 }
}
]
}
]
}
}
Program builder
The ProgramBuilder class manages instruction accumulation:
class ProgramBuilder {
private instructions: Program.Instruction[] = [];
addInstruction(
offset: number,
operation: Program.Instruction.Operation,
context?: Program.Context,
) {
this.instructions.push({ offset, operation, context });
}
build(): Program {
return {
instructions: this.instructions,
};
}
}
Variable scoping
BUG includes storage variables in the variables context for every instruction that accesses them. This allows debuggers to inspect storage values at any execution point.
export function collectVariablesWithLocations(
state: State,
sourceId: string,
): VariableInfo[] {
const variables: VariableInfo[] = [];
// Storage variables have fixed slots
for (const storageDecl of state.module.storageDeclarations) {
const bugType = state.types.get(storageDecl.id);
const pointer = generateStoragePointer(storageDecl.slot, bugType);
variables.push({
identifier: storageDecl.name,
type: convertBugType(bugType),
pointer,
declaration: storageDecl.loc
? {
source: { id: sourceId },
range: storageDecl.loc,
}
: undefined,
});
}
return variables;
}
BUG also lists the local variables (parameters and lets) in scope at each
instruction, each with a pointer to its value where the value is located: a
memory word in the function's frame, or a stack slot. A local hides a storage
variable of the same name while it is in scope. Under optimization, a local
whose value the optimizer folded to a constant or removed is listed without a
pointer. See packages/bugc/src/evmgen/debug/local-variables.ts.
Testing strategy
BUG tests debug output in several ways:
- Unit tests — Test individual conversion functions
- Integration tests — Compile programs and verify output structure
- Playground — Visual verification of output (see BUG playground)
Example test pattern:
test("generates correct storage pointer", () => {
const source = `
name Test;
storage { [0] value: uint256; }
code { value = 42; }
`;
const result = compile({ to: "bytecode", source });
const program = result.value.program;
// Find the SLOAD instruction
const sload = program.instructions.find(
(i) => i.operation?.mnemonic === "SLOAD",
);
expect(sload.context).toMatchObject({
variables: [
{
identifier: "value",
pointer: { location: "storage", slot: 0 },
},
],
});
});
Lessons learned
Start simple
BUG started with just storage variables and elementary types. Complex features (arrays, mappings, structs) were added incrementally after the basic infrastructure was working.
Preserve information early
Type information is easier to preserve than reconstruct. BUG's IR types carry their source-level origin, which simplifies later conversion.
Test visually
The BUG playground proved invaluable for debugging the debug output. Being able to see the generated annotations alongside bytecode helped catch issues that unit tests missed.
Handle edge cases gracefully
When storage analysis can't determine a location (e.g., computed slot from a non-constant), BUG still generates useful partial information rather than failing entirely.
Future work
Areas for improvement in BUG's debug support:
- Richer contexts — Add frame information for function calls
- Source maps — More granular source location tracking
Resources
- BUG playground — Interactive compiler demo
- BUG Source Code — Implementation reference
- ethdebug/format Specification — Format reference