{"id":595,"date":"2024-09-14T13:45:59","date_gmt":"2024-09-14T10:45:59","guid":{"rendered":"https:\/\/netbsd.name\/?p=595"},"modified":"2024-09-14T13:45:59","modified_gmt":"2024-09-14T10:45:59","slug":"resolving-memory-corruption-in-bios-bootloader-on-via-c7-based-systems","status":"publish","type":"post","link":"https:\/\/netbsd.name\/?p=595","title":{"rendered":"Resolving Memory Corruption in BIOS bootloader on VIA C7-Based Systems"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Following the resolution of the ALTINST mode issue, I returned to investigating the VIA C7 boot problem. Under certain circumstances, some Esther-based systems were experiencing a sudden reboot shortly after the kernel was loaded, which I was able to reproduce on my <a href=\"https:\/\/www.biostar.com.tw\/app\/en\/eol\/introduction.php?S_ID=470\">Biostar Viotech 3100+<\/a> motherboard. This led to a lengthy debugging process before I could finally identify the culprit!<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Initial analysis<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Typically, the bugs I encountered before were related to the NetBSD kernel or drivers. Initially, I assumed that this might also be an issue with the INSTALL kernel configuration because, during the early stages of investigation, the problem only occurred with the install image (not with the fully installed system on my SD card). A critical discovery in the debugging process was that the issue specifically occurred when ACPI 3.0 was enabled. I downloaded older releases to identify the first affected version and found that NetBSD 7.0 was the first to exhibit the symptoms.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To narrow down the problematic commit, I decided to install an older release on my USB image and build several kernels between different 6.99.x versions. However, I soon realized that the kernel wasn&#8217;t the issue &#8211; older 6.x kernels were also failing with the newer install images. At the same time, the install kernel was successfully booting from my SD card setup. At this point, it became clear that the problem was with the bootloader.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I then began building full distributions from various 6.99.x versions to pinpoint the commit responsible. This process was slow and painful, taking between 5 to 8 hours for each build. In hindsight, I could have just been building the bootloader code, but at the time, I didn&#8217;t know where the affected code was or how to build it. After several weeks of this repetitive process, I finally identified that the reboot issue began after the switch to GCC 4.8, specifically when the boot parameters were fixed in this particular <a href=\"https:\/\/github.com\/NetBSD\/src\/commit\/ad33dd774c2fe8beb41c96d1d29aef4ebce3f5cb\">commit<\/a>. Unfortunately, this didn&#8217;t offer much insight into the underlying problem, and this approach reached a dead end.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">At that point, I returned to the current code and started debugging the kernel&#8217;s behavior.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Kernel debugging<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The boot log was printing only a few messages before the reboot, but it still provided a useful starting point, especially the last line: &#8216;pmap_kenter_pa: mapping already present&#8217;. I quickly located this message in the <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/a4816d71d08f17e4b301a30d01947d54dde2ded7\/sys\/arch\/x86\/x86\/pmap.c#L1054\">code<\/a> and began investigating what was happening. The comment in the conditional block stated, &#8216;This should not happen,&#8217; but clearly, it was. Eventually, the code called the <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/a4816d71d08f17e4b301a30d01947d54dde2ded7\/sys\/arch\/x86\/x86\/x86_tlb.c#L289\"><em>kcpuset_copy()<\/em><\/a> method, where both arguments were still undefined, leading to a null pointer dereference during the <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/a4816d71d08f17e4b301a30d01947d54dde2ded7\/sys\/kern\/subr_kcpuset.c#L351\"><em>memcpy()<\/em><\/a> call and triggering a sudden reboot.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Knowing this was helpful, but it didn\u2019t explain why this &#8216;should not happen&#8217; situation was occurring. Due to the early stage of the kernel boot process, getting a useful stack trace was either difficult or impossible. Nevertheless, I began tracing the <em>pmap_kenter_pa()<\/em> calls and hypothesizing where the call could have originated, especially since I knew that the global <em>kcpuset_running<\/em> parameter was not supposed to be set at this point. This is where comparing the behavior of ACPI 2.0 and ACPI 3.0 became useful. I inserted various debugging messages in the relevant parts of the code and compared the memory values between the two. This quickly led me to discover invalid virtual memory values that were dependent on the parameters passed from the bootloader (e.g., atdevbase, PDPpaddr).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It took some time to pinpoint exactly where the problematic code was being executed. My initial assumption was that the issue occurred somewhere in kern\/init_main.c <em><a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/a4816d71d08f17e4b301a30d01947d54dde2ded7\/sys\/kern\/init_main.c#L263\">main()<\/a>,<\/em> but it turned out to be earlier, in <em>init386()<\/em> (specifically <em><a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/a4816d71d08f17e4b301a30d01947d54dde2ded7\/sys\/arch\/i386\/i386\/machdep.c#L1343\">init386_pte0()<\/a><\/em>), which was making the problematic <em><a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/a4816d71d08f17e4b301a30d01947d54dde2ded7\/sys\/arch\/i386\/i386\/machdep.c#L1081C2-L1081C16\">pmap_kenter_pa()<\/a><\/em> call. At this stage, <em>kcpuset_running<\/em> is not yet initialized, as that occurs later. However, since the conditional block causing the reboot wasn\u2019t supposed to be executed so early in the process, no assertions were added to that part of the code.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Despite this progress, I was slowly hitting a dead end again. It was clear that something was wrong with the memory values, and code inspection showed their dependency on the bootloader\u2019s input, potentially causing a reboot when these values deviated too far from the expected range. This also explained why the boot process sometimes succeeded. I could identify some workarounds at this point\u2014such as ignoring the <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/a4816d71d08f17e4b301a30d01947d54dde2ded7\/sys\/arch\/i386\/i386\/locore.S#L858\">eblob<\/a> value in the calculations\u2014but these were not viable long-term solutions. It was becoming clear that I needed to start debugging the bootloader itself!<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">BIOS bootloader<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">At this point in the analysis, I already knew that the affected bootloader code was located in the <em>sys\/arch\/i386\/stand<\/em> path and that it was part of the biosboot bootloader. With the help of other NetBSD developers and the documentation, I learned how to build the bootloader alone and install it into my installation image. This greatly sped up the debugging process, as I no longer needed to build the full distribution.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This process is relatively simple using the <a href=\"https:\/\/www.netbsd.org\/docs\/guide\/en\/chap-build.html#chap-boot-cross-build-kernel\">build.sh<\/a> framework:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><code># build i386 cross-compile toolchain\n.\/build.sh -T ..\/tools -O ..\/obj -U -j6 -mi386 tools\n# to avoid searching for all dependencies repeatedly, build the distribution once using the build.sh framework\n.\/build.sh -T ..\/tools -O ..\/obj -U -j6 -mi386 distribution\n# navigate to the i386 bootloaders code folder\ncd sys\/arch\/i386\/stand\/\n# build the bootloaders and repeat the process as many times as necessary\n..\/tools\/bin\/nbmake-i386 -j6 dependall\n# install to destdir\n..\/tools\/bin\/nbmake-i386 -j6 install<\/code><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Then install the bios bootloader:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><code># mount NetBSD install image\nmount \/dev\/sd0a \/mnt\n# copy secondary bootstrap to the root folder\nsudo cp ..\/obj\/destdir.i386\/usr\/mdec\/boot \/mnt\/boot\n# copy bootxx_* files (likely optional)\nsudo cp ..\/obj\/destdir.i386\/usr\/mdec\/* \/mnt\/usr\/mdec\/\n# install the primary bootstrap\ninstallboot \/dev\/sd0a ..\/obj\/destdir.i386\/usr\/mdec\/bootxx_ffsv1\n# unmount install image\numount \/mnt<\/code><\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">To avoid constantly mounting, unmounting, and re-attaching the USB stick, files can also be transferred via SSH and installed directly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The main logic of the bootloader resides in the <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/6622bba852cc4518c5b65e975373482d20dc4a20\/sys\/arch\/i386\/stand\/lib\/exec.c#L412C1-L412C12\"><em>exec_netbsd()<\/em><\/a> function, which is called by various i386 bootloaders with parameters from the primary bootloader&#8217;s input. This function loads the kernel, calculates its size, and performs related tasks. My debugging process focused on identifying where memory values were becoming incorrect. The <em>marks[]<\/em> array, where values are set during the kernel load process, was a major point of focus.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">After multiple attempts, I discovered that the initial values in <em>marks[]<\/em> were correct immediately after the kernel was <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/6622bba852cc4518c5b65e975373482d20dc4a20\/sys\/arch\/i386\/stand\/lib\/exec.c#L383\">loaded<\/a>, but they became corrupted by the end of the <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/6622bba852cc4518c5b65e975373482d20dc4a20\/sys\/arch\/i386\/stand\/lib\/exec.c#L405\"><em>common_load_kernel()<\/em><\/a> method. The corruption occurred despite only a few calls happening between these points, making it easier to identify the cause. The <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/6622bba852cc4518c5b65e975373482d20dc4a20\/sys\/arch\/i386\/stand\/lib\/exec.c#L400C2-L400C17\"><em>bi_getmemmap()<\/em><\/a> call was pinpointed as the culprit. It was identified that the stack overflow occurred upon returning from this method, leading to stack corruption.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The issue was eventually narrowed down to the <a href=\"https:\/\/github.com\/NetBSD\/src\/blob\/6622bba852cc4518c5b65e975373482d20dc4a20\/sys\/arch\/i386\/stand\/lib\/biosmemx.S#L112\"><em>getmementry()<\/em><\/a> call, which is invoked multiple times within <em>bi_getmemmap<\/em>(). Even a single call to <em>getmementry()<\/em> was enough to corrupt the stack right after returning from <em>bi_getmemmap()<\/em>. Debugging this was challenging because <em>getmementry()<\/em> is written in assembly code and has not been modified for many years.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">With assistance, I eventually found that the allocated buffer for 5 words was actually writing to 6 words when ACPI 3.0 was enabled. It appeared that ACPI 3.0 extended the <a href=\"https:\/\/wiki.osdev.org\/Detecting_Memory_(x86)#BIOS_Function:_INT_0x15,_EAX_=_0xE820\">INT 0x15, EAX = 0xE820<\/a> BIOS function for memory detection from 20 bytes to 24 bytes to accommodate extended attributes. Only a few motherboards supported 24 bytes initially, while the specification required that the function return 20 bytes if requested, regardless of the actual support for 24 bytes. Some VIA systems shipped with a buggy BIOS that returned 24 bytes regardless.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The temporary buffer was not allocated for 24 bytes, causing a stack buffer overrun. The <a href=\"https:\/\/github.com\/NetBSD\/src\/commit\/90d347b4b6c7226624933877a98193c88b7054e3\">fix<\/a> was to increase the buffer size from 5 to 6 words!<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Conclusion<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The Biostar Viotech 3100+ now boots successfully, with memory values consistent whether ACPI 3.0 is enabled or not. Changes have been applied to the NetBSD 10 and NetBSD 9 branches (as older releases are no longer supported). I must admit, this investigation was challenging: it involved a long process of narrowing down the issue, making incorrect and time-consuming decisions, countless reboots, and numerous bootloader builds and reinstalls. It required many long evenings to make slow progress or to rule out incorrect theories.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One might question whether spending time on outdated systems is worthwhile, and the answer might be no. However, like every solved mystery, it provides significant rewards in terms of knowledge and experience. I believe this exercise was valuable for me, and if even one user benefits from this fix, it will have been worth it. This motherboard has allowed me to address several issues: from a broken temperature sensor and a faulty ATLINST mode disable process to finally resolving the boot process failure. Now, I can give it some well-deserved rest and shift my focus to other tasks.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Following the resolution of the ALTINST mode issue, I returned to investigating the VIA C7 boot problem. Under certain circumstances, some Esther-based systems were experiencing a sudden reboot shortly after the kernel was loaded, which I was able to reproduce on my Biostar Viotech 3100+ motherboard. This led to a lengthy debugging process before I &hellip; <a href=\"https:\/\/netbsd.name\/?p=595\" class=\"more-link\">Continue reading <span class=\"screen-reader-text\">Resolving Memory Corruption in BIOS bootloader on VIA C7-Based Systems<\/span> <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-595","post","type-post","status-publish","format-standard","hentry","category-netbsd"],"_links":{"self":[{"href":"https:\/\/netbsd.name\/index.php?rest_route=\/wp\/v2\/posts\/595","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/netbsd.name\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/netbsd.name\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/netbsd.name\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/netbsd.name\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=595"}],"version-history":[{"count":9,"href":"https:\/\/netbsd.name\/index.php?rest_route=\/wp\/v2\/posts\/595\/revisions"}],"predecessor-version":[{"id":605,"href":"https:\/\/netbsd.name\/index.php?rest_route=\/wp\/v2\/posts\/595\/revisions\/605"}],"wp:attachment":[{"href":"https:\/\/netbsd.name\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=595"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/netbsd.name\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=595"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/netbsd.name\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=595"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}