incremental compilation: reached unreachable code and accumulation of errors if `@import`ing non-existent files #22696

llogick · 2025-01-31T17:13:34Z

Zig Version

0.14.0-dev.3020+c104e8644 DEBUG build

Steps to Reproduce and Observed Behavior

Start out with a zig init project
Add a "check" step for the exe
build with zig build -fincremental --watch check
Add an @import to a non existent file in src/main.zig, eg const argh = @import("blah.zig");
Save main.zig (compiler will report src/main.zig:x:y: error: unable to load 'src/blah.zig': FileNotFound)
Modify the @import string, eg delete the 'g' -> const argh = @import("blah.zi");
Save main.zig

zig build -fincremental --watch check
Build Summary: 2/2 steps succeeded
check success
└─ zig build-exe zig-init Debug native success 2s
check
└─ zig build-exe zig-init Debug native 1 errors
src/main.zig:5:22: error: unable to load 'src/blah.zig': FileNotFound
const argh = @import("blah.zig");
                     ^~~~~~~~~~
error: 1 compilation errors
Build Summary: 0/2 steps succeeded; 1 failed
check transitive failure
└─ zig build-exe zig-init Debug native 1 errors
check
└─ zig build-exe zig-init Debug native failure
error: thread 413716 panic: reached unreachable code
/home/rad/lab/zig/build/stage4/lib/zig/std/posix.zig:1820:23: 0x6063fdf in openatZ (zig)
            .NOENT => return error.FileNotFound,
                      ^
/home/rad/lab/zig/build/stage4/lib/zig/std/fs/Dir.zig:886:16: 0x5ed8d12 in openFileZ (zig)
    const fd = try posix.openatZ(self.fd, sub_path, os_flags, 0);
               ^
/home/rad/lab/zig/build/stage4/lib/zig/std/fs/Dir.zig:833:5: 0x5d3db85 in openFile (zig)
    return self.openFileZ(&path_c, flags);
    ^
/home/rad/lab/zig/build/stage4/lib/zig/std/Build/Cache/Path.zig:62:5: 0x5dea14d in openFile (zig)
    return p.root_dir.handle.openFile(joined_path, flags);
    ^
/home/rad/lab/zig/src/Zcu/PerThread.zig:92:23: 0x66d5013 in astGenFile (zig)
    var source_file = try file.mod.root.openFile(file.sub_file_path, .{});
                      ^
/home/rad/lab/zig/build/stage4/lib/zig/std/debug.zig:518:14: 0x5d37eed in assert (zig)
    if (!ok) unreachable; // assertion failure
             ^
/home/rad/lab/zig/build/stage4/lib/zig/std/array_hash_map.zig:966:19: 0x62ff86d in putNoClobberContext (zig)
            assert(!result.found_existing);
                  ^
/home/rad/lab/zig/build/stage4/lib/zig/std/array_hash_map.zig:962:44: 0x60d7cd1 in putNoClobber (zig)
            return self.putNoClobberContext(gpa, key, value, undefined);
                                           ^
/home/rad/lab/zig/src/Zcu/PerThread.zig:3241:42: 0x66d75f4 in reportRetryableAstGenError (zig)
        try zcu.failed_files.putNoClobber(gpa, file, err_msg);
                                         ^
/home/rad/lab/zig/src/Compilation.zig:4284:42: 0x62fb2c2 in workerAstGenFile (zig)
            pt.reportRetryableAstGenError(src, file_index, err) catch |oom| switch (oom) {
                                         ^
/home/rad/lab/zig/build/stage4/lib/zig/std/Thread/Pool.zig:182:50: 0x62fb9ee in runFn (zig)
            @call(.auto, func, .{id.?} ++ closure.arguments);
                                                 ^
/home/rad/lab/zig/build/stage4/lib/zig/std/Thread/Pool.zig:295:32: 0x62a9364 in worker (zig)
            run_node.data.runFn(&run_node.data, id);
                               ^
/home/rad/lab/zig/build/stage4/lib/zig/std/Thread.zig:488:13: 0x608cc2a in callFn__anon_175265 (zig)
            @call(.auto, f, args);
            ^
/home/rad/lab/zig/build/stage4/lib/zig/std/Thread.zig:757:30: 0x5f14c24 in entryFn (zig)
                return callFn(f, args_ptr.*);
                             ^
???:?:?: 0x7f10e6098291 in ??? (libc.so.6)
Unwind information for `libc.so.6:0x7f10e6098291` was not available, trace may be incomplete


error: the following command terminated unexpectedly:
/home/rad/lab/zig/build/stage5/bin/zig build-exe -ODebug -Mroot=/home/rad/lab/zta/misc/zig-init/src/main.zig -fno-emit-bin --cache-dir /home/rad/lab/zta/misc/zig-init/.zig-cache --global-cache-dir /home/rad/.cache/zig --name zig-init --zig-lib-dir /home/rad/lab/zig/build/stage5/lib/zig/ -fincremental --listen=-

A (Zig) Release build will just continue to report (even if you comment out the import statement)

check
└─ zig build-exe zig-init Debug native 1 errors
error: unable to load 'src/blah.zig': FileNotFound
error: 1 compilation errors
Build Summary: 0/2 steps succeeded; 1 failed

Modifying the import statement causes accumulation of errors to nonexistent files

check transitive failure
└─ zig build-exe zig-init Debug native 2 errors
check
└─ zig build-exe zig-init Debug native 3 errors
error: unable to load 'src/blah.zig': FileNotFound
error: unable to load 'src/bleh.zig': FileNotFound
src/main.zig:5:22: error: unable to load 'src/hmm.zig': FileNotFound
const argh = @import("hmm.zig");
                     ^~~~~~~~~
error: 3 compilation errors
Build Summary: 0/2 steps succeeded; 1 failed
check transitive failure
└─ zig build-exe zig-init Debug native 3 errors
check
└─ zig build-exe zig-init Debug native 3 errors
error: unable to load 'src/blah.zig': FileNotFound
error: unable to load 'src/bleh.zig': FileNotFound
error: unable to load 'src/hmm.zig': FileNotFound
error: 3 compilation errors

Expected Behavior

The text was updated successfully, but these errors were encountered:

mlugg · 2025-02-04T15:17:37Z

Oh, this one's a bit annoying. After running the AstGen workers at the start of an update, we'll need to do a quick traversal of files to check which ones are actually imported anywhere.

From there, we can set a flag on files to avoid emitting errors for unreferenced ones, and (if we proceed to analysis) lose the ZIR instructions associated with unreferenced files (since they can't be allowed to participate in this compilation).

This commit makes some big changes to how we track state for Zig source files. In particular, it changes: * How `File` tracks its path on-disk * How AstGen discovers files * How file-level errors are tracked * How `builtin.zig` files and modules are created There is also one breaking change here, which is that `@import` of a non-existent module is now a compile error even if the import is not semantically analyzed. The original motivation here was to address incremental compilation bugs with the handling of files, such as ziglang#22696. To fix this, a few changes are necessary. Just like declarations may become unreferenced on an incremental update, meaning we suppress analysis errors associated with them, it is also possible for all imports of a file to be removed on an incremental update, in which case file-level errors for that file should be suppressed. As such, after AstGen, the compiler must traverse files (starting from analysis roots) and discover the set of "live files" for this update. Additionally, the compiler's previous handling of retryable file errors was not very good; the source location the error was reported as was based only on the first discovered import of that file. This source location also disappeared on future incremental updates. So, as a part of the file traversal above, we also need to figure out the source locations of imports which errors should be reported against. Another observation I made is that the "file exists in multiple modules" error was not implemented in a particularly good way (I get to say that because I wrote it!). It was subject to races, where the order in which different imports of a file were discovered affects both how errors are printed, and which module the file is arbitrarily assigned, with the latter in turn affecting which other files are considered for import. The thing I realised here is that while the AstGen worker pool is running, we cannot know for sure which module(s) a file is in; we could always discover an import later which changes the answer. So, here's how the AstGen workers have changed. We initially ensure that `zcu.import_table` contains the root files for all modules in this Zcu, even if we don't know any imports for them yet. Then, the AstGen workers do not need to be aware of modules. Instead, they simply ignore module imports, and only spin off more workers when they see a by-path import. During this phase, we work exclusively in absolute paths, to ensure consistency in terms of `NameTooLong` errors. After the AstGen workers all complete, we know that any file which might be imported is definitely in `import_table` and up-to-date. So, we perform a single-threaded graph traversal; similar to what `resolveReferences` plays for `AnalUnit`s, but for files instead. We figure out which files are alive, and which module each file is in. If a file turns out to be in multiple modules, we set a field on `Zcu` to indicate this error. If a file is in a different module to a prior update, we set a flag instructing `updateZirRefs` to invalidate all dependencies on the file. This traversal also discovers "import errors"; these are errors associated with a specific `@import`. There are two possible errors here: "module not found" when a module import uses an unmapped name, or "import outside of module root" when importing a file using `..` past the module root path. These errors have to be identified during this traversal because they depend on which module the file is in, which we don't know until now. For simplicity, `failed_files` now just maps to `?[]u8`, since the source location is always the whole file. In fact, this allows removing `LazySrcLoc.Offset.entire_file` completely, slightly simplifying some error reporting logic. File-level errors are now directly built in the `std.zig.ErrorBundle.Wip`. If the payload is not `null`, it is the message for a retryable error (i.e. an error loading the source file), and will be reported with a "file imported here" note pointing to the import site discovered during the single-threaded file traversal. The last piece of fallout here is how `Builtin` works. Rather than constructing "builtin" modules when creating `Package.Module`s, they are now constructed on-the-fly by `Zcu`. The map `Zcu.builtin_modules` maps from digests to `*Package.Module`s. These digests are abstract hashes of the `Builtin` value; i.e. all of the options which are placed into "builtin.zig". During the file traversal, we populate `builtin_modules` as needed, so that when we see this imports in Sema, we just grab the relevant entry from this map. This eliminates a bunch of awkward state tracking during construction of the module graph. It's also now clearer exactly what options the builtin module has, since previously it inherited some options arbitrarily from the first-created module with that "builtin" module! The actual effects of this commit are: * retryable file errors are now consistently reported against the whole file, with a note pointing to a live import of that file * incremental updates do not print retryable file errors differently between updates * incremental updates support files changing modules * incremental updates support files becoming unreferenced Resolves: ziglang#22696

This commit makes some big changes to how we track state for Zig source files. In particular, it changes: * How `File` tracks its path on-disk * How AstGen discovers files * How file-level errors are tracked * How `builtin.zig` files and modules are created The original motivation here was to address incremental compilation bugs with the handling of files, such as ziglang#22696. To fix this, a few changes are necessary. Just like declarations may become unreferenced on an incremental update, meaning we suppress analysis errors associated with them, it is also possible for all imports of a file to be removed on an incremental update, in which case file-level errors for that file should be suppressed. As such, after AstGen, the compiler must traverse files (starting from analysis roots) and discover the set of "live files" for this update. Additionally, the compiler's previous handling of retryable file errors was not very good; the source location the error was reported as was based only on the first discovered import of that file. This source location also disappeared on future incremental updates. So, as a part of the file traversal above, we also need to figure out the source locations of imports which errors should be reported against. Another observation I made is that the "file exists in multiple modules" error was not implemented in a particularly good way (I get to say that because I wrote it!). It was subject to races, where the order in which different imports of a file were discovered affects both how errors are printed, and which module the file is arbitrarily assigned, with the latter in turn affecting which other files are considered for import. The thing I realised here is that while the AstGen worker pool is running, we cannot know for sure which module(s) a file is in; we could always discover an import later which changes the answer. So, here's how the AstGen workers have changed. We initially ensure that `zcu.import_table` contains the root files for all modules in this Zcu, even if we don't know any imports for them yet. Then, the AstGen workers do not need to be aware of modules. Instead, they simply ignore module imports, and only spin off more workers when they see a by-path import. During AstGen, we can't use module-root-relative paths, since we don't know which modules files are in; but we don't want to unnecessarily use absolute files either, because those are non-portable and can make `error.NameTooLong` more likely. As such, I have introduced a new abstraction, `Compilation.Path`. This type is a way of representing a filesystem path which has a *canonical form*. The path is represented relative to one of a few special directories: the lib directory, the global cache directory, or the local cache directory. As a fallback, we use absolute (or cwd-relative on WASI) paths. This is kind of similar to `std.Build.Cache.Path` with a pre-defined list of possible `std.Build.Cache.Directory`, but has stricter canonicalization rules based on path resolution to make sure deduplicating files works properly. A `Compilation.Path` can be trivially converted to a `std.Build.Cache.Path` from a `Compilation`, but is smaller, has a canonical form, and has a digest which will be consistent across different compiler processes with the same lib and cache directories (important when we serialize incremental compilation state in the future). `Zcu.File` and `Zcu.EmbedFile` both contain a `Compilation.Path`, which is used to access the file on-disk; module-relative sub paths are used quite rarely (`EmbedFile` doesn't even have one now for simplicity). After the AstGen workers all complete, we know that any file which might be imported is definitely in `import_table` and up-to-date. So, we perform a single-threaded graph traversal; similar to what `resolveReferences` plays for `AnalUnit`s, but for files instead. We figure out which files are alive, and which module each file is in. If a file turns out to be in multiple modules, we set a field on `Zcu` to indicate this error. If a file is in a different module to a prior update, we set a flag instructing `updateZirRefs` to invalidate all dependencies on the file. This traversal also discovers "import errors"; these are errors associated with a specific `@import`. With Zig's current design, there is only one possible error here: "import outside of module root". This must be identified during this traversal instead of during AstGen, because it depends on which module the file is in. I tried also representing "module not found" errors in this same way, but it turns out to be much more useful to report those in Sema, because of use cases like optional dependencies where a module import is behind a comptime-known build option. For simplicity, `failed_files` now just maps to `?[]u8`, since the source location is always the whole file. In fact, this allows removing `LazySrcLoc.Offset.entire_file` completely, slightly simplifying some error reporting logic. File-level errors are now directly built in the `std.zig.ErrorBundle.Wip`. If the payload is not `null`, it is the message for a retryable error (i.e. an error loading the source file), and will be reported with a "file imported here" note pointing to the import site discovered during the single-threaded file traversal. The last piece of fallout here is how `Builtin` works. Rather than constructing "builtin" modules when creating `Package.Module`s, they are now constructed on-the-fly by `Zcu`. The map `Zcu.builtin_modules` maps from digests to `*Package.Module`s. These digests are abstract hashes of the `Builtin` value; i.e. all of the options which are placed into "builtin.zig". During the file traversal, we populate `builtin_modules` as needed, so that when we see this imports in Sema, we just grab the relevant entry from this map. This eliminates a bunch of awkward state tracking during construction of the module graph. It's also now clearer exactly what options the builtin module has, since previously it inherited some options arbitrarily from the first-created module with that "builtin" module! The user-visible effects of this commit are: * retryable file errors are now consistently reported against the whole file, with a note pointing to a live import of that file * some theoretical bugs where imports are wrongly considered distinct (when the import path moves out of the cwd and then back in) are fixed * some consistency issues with how file-level errors are reported are fixed; these errors will now always be printed in the same order regardless of how the AstGen pass assigns file indices * incremental updates do not print retryable file errors differently between updates or depending on file structure/contents * incremental updates support files changing modules * incremental updates support files becoming unreferenced Resolves: ziglang#22696

llogick added the bug Observed behavior contradicts documented or intended behavior label Jan 31, 2025

mlugg added the incremental compilation Problem occurs only when reusing compiler state. label Jan 31, 2025

llogick mentioned this issue Feb 1, 2025

incremental compilation: reached unreachable code and accumulation of errors if @importing non-existent files llogick/zigscient-next#3

Open

mlugg added this to the 0.14.0 milestone Feb 4, 2025

andrewrk modified the milestones: 0.14.0, 0.14.1 Mar 1, 2025

andrewrk modified the milestones: 0.14.1, 0.15.0 Apr 15, 2025

mlugg closed this as completed in 37a9a4e May 20, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Uh oh!

incremental compilation: reached unreachable code and accumulation of errors if `@import`ing non-existent files #22696

incremental compilation: reached unreachable code and accumulation of errors if `@import`ing non-existent files #22696

llogick commented Jan 31, 2025

mlugg commented Feb 4, 2025

Uh oh!

Uh oh!

incremental compilation: reached unreachable code and accumulation of errors if @importing non-existent files #22696

incremental compilation: reached unreachable code and accumulation of errors if @importing non-existent files #22696

Comments

llogick commented Jan 31, 2025

Zig Version

Steps to Reproduce and Observed Behavior

Expected Behavior

mlugg commented Feb 4, 2025

Uh oh!

incremental compilation: reached unreachable code and accumulation of errors if `@import`ing non-existent files #22696

incremental compilation: reached unreachable code and accumulation of errors if `@import`ing non-existent files #22696