# EXPERIMENTAL This library is still in an experimental stage and should not be used in environments where stability/safety/security/reliability are important. Some minor features like normalization and recursive entity references aren't supported yet. While this library is in an experimental state, semantic versioning conventions will not be guaranteed and minor updates may introduce breaking API/ABI changes. Semantic versioning will be guaranteed once this library hits 1.0. # Why use imxml? imxml is a portable, high-performance, zero-allocation, immediate-mode XML parser designed for applications that know the approximate structure of the XML they consume. The rationale of this library is that you generally don't need to parse arbitrary XML. Instead, you're targeting some subset. This goes beyond just disabling certain features, as even the structure of your XML is typically bounded. For example, if you're parsing OpenStreetMap (OSM) data, you know that you'll have one root element, which contains one element, followed by several elements, which may contain elements, then several elements, and so on. This "immediate-mode" API design allows you to skip parsing anything you're not interested in, while also indicating your data's expected layout to the compiler, which allows the optimizer to do a better job. ## Performance benchmarks/imxml/benchmark-results.xml contains benchmark results for some workloads on an AMD Ryzen 5 3600. In these benchmarks, imxml achieves a minimum 2.9x performance improvement compared to quick-xml (the fastest mainstream parser I could find) for these workloads, and up to 12x in extreme cases. Note that these are synthetic benchmarks designed to stress the raw parsing throughput and nothing else. Your code likely contains all sorts of extra business logic that will change your perceived throughput, though imxml's immediate-mode API design should give your optimizer the best chance to generate fast code. ## Memory imxml's memory usage is completely static, there are no allocations, and it doesn't modify the input buffer. Returned strings are simply views into the input buffer or, in the case of character references, views into a static (optionally thread-local) buffer. ## Compatibility Standard C89 (ANSI C) and C++98 are supported. The library has been tested with clang, gcc, and tcc. Every standards-compliant C compiler should work, though if emmintrin.h/immintrin.h isn't supported, only the fallback scalar implementation will work (unless you define your own instruction set). UTF-8, ISO-Latin-1, ASCII, and generally all char/byte-based encodings are supported. UTF-16 is not supported. The library itself has no dependencies, not even libc, but optional platform-specific convenience functions can be enabled with flags. Right now, there is only a Linux helper implementation for loading files. ## Extensibility Currently imxml supports the MMX, SSE2, and AVX2 instruction sets, as well as a fallback scalar configuration for compilers that don't support SIMD intrinsics (e.g. tcc). You can define custom instruction set definitions if you want SIMD acceleration on more exotic platforms. ## Features - [X] Entities - [X] Standard entity references - [X] Unicode character references - [X] DTD-defined entity references - [X] CDATA - [X] Comments - [X] Namespaces - [X] Customizable SIMD instruction sets - [ ] Content normalization - [ ] Platform-specific helper functions - [X] Linux - [ ] Windows - [ ] OSX # Usage Download imxml.h and add this to your code: ``` #define IMXML_IMPLEMENTATION #include "imxml.h" ``` ## Example ``` #define IMXML_LINUX // We're using Linux, of course #define IMXML_NO_SUPPORT_SINGLE_QUOTES // OSM always uses double-quotes #define IMXML_NO_SUPPORT_COMMENTS // OSM doesn't include comments #define IMXML_NO_SUPPORT_CDATA // OSM doesn't include CDATA #define IMXML_NO_SUPPORT_ENTITIES // We don't need to parse entities #define IMXML_NO_SUPPORT_NAMESPACES // OSM doesn't include namespaces #define IMXML_NO_CHECK_BOUNDS // We know our data will end with #define THREADLOCAL // We disable thread-local by defining it as nothing #define IMXML_IMPLEMENTATION #include "imxml.h" #include // for printf int main (void) { size_t file_size; // We need to use this file_open function because the input data needs to // be padded to make sure we're not reading out-of-bounds char* const file = imxml_linux_file_open("dat/yellowstone.osm", &file_size, false); if (file == 0) return -1; XmlParser p = {0}; // zero-init is important p.head = file; // Verify we're really dealing with osm data. This is mostly unnecessary, // but it doesn't cost much and might catch some dumb bugs if (!xml_parse_header(&p, NULL)) return -1; if (!xml_tag_expect(&p, "osm")) return -1; if (!xml_has_children(&p)) return -1; if (!xml_tag_expect(&p, "bounds")) return -1; if (xml_has_children(&p)) return -1; size_t highway_crossing_count = 0; while (true) { ImxmlString tag = xml_tag(&p); if (imxml_streql(tag, imxml_strlit("node"))) { if (xml_has_children(&p)) { while (true) { tag = xml_tag(&p); if (imxml_streql(tag, imxml_strlit("/node"))) break; ImxmlString key = xml_value(&p); if (!imxml_streql(key, imxml_strlit("highway"))) continue; ImxmlString value = xml_value(&p); if (imxml_streql(value, imxml_strlit("crossing"))) { highway_crossing_count += 1; } } } } else { // encountered non-node tag, which means we can early-out (because // OSM groups all of the elements up-front). // Alternatively, we could check if tag == "/osm", but then we'd // just be skipping over all the and tags, which // would be a waste of time break; } } printf("%zu\n", highway_crossing_count); imxml_linux_file_close(file, file_size); return 0; } ``` See the examples/ and benchmarks/imxml/ directories for more examples. # Contributing If you have any ideas or issues, send them to steven@vandorp.lu. If you want to modify code, make your changes in a separate branch, generate the patch(es) (`git format-patch master` from inside your branch) and send the patch file(s) to steven@vandorp.lu. # Documentation See the DOCUMENTATION section in the source code.