AngleSharp HTML5 Spec Compliance: mXSS via annotation-xml HTML Integration Point Bypass
Description
Summary
The HTML specification requires that a MathML ` element with encoding="text/html" or encoding="application/xhtml+xml"` is treated as an HTML integration point. Content inside it must be parsed as HTML, not MathML.
AngleSharp does not implement this correctly. As a result, the parser produces a DOM tree that differs from what a browser will build (different namespaces if encoding="text/html" is not treated) when given the same serialized output. Two bugs combine to make this exploitable:
- Missing HtmlTip flag: MathAnnotationXmlElement is never assigned NodeFlags.HtmlTip based on its encoding attribute, so the Consume() dispatch always routes tokens to Foreign() instead of Home() (HTML mode). - Unescaped < > in attribute values: HtmlMarkupFormatter.WriteAttributeValue() does not escape < or > characters, only & and ". This allows injected markup to break out of attribute values on re-parse. _See Escape "<" and ">" in attributes when serializing HTML #6235 _
Details
In MathAnnotationXmlElement (AngleSharp/Mathml/Dom/Internal/MathAnnotationXmlElement.cs): ``cs // Current — HtmlTip is never set : base(owner, TagNames.AnnotationXml, prefix, NodeFlags.Special | NodeFlags.Scoped) ``
Because HtmlTip is absent, the token dispatch in Consume() always sends tokens to Foreign() when inside annotation-xml, regardless of the encoding attribute. The compensating check in ForeignNormalTag() only covers tags in AllForeignExceptions and is entirely bypassed during fragment parsing (innerHTML setter) due to an if (!IsFragmentCase) guard.
In HtmlMarkupFormatter.WriteAttributeValue() (AngleSharp/Html/HtmlMarkupFormatter.cs): ``cs // Escapes & " and \u00A0, but NOT < or > case Symbols.Ampersand: stringBuilder.Append("&"); break; case Symbols.NoBreakSpace: stringBuilder.Append(" "); break; case Symbols.DoubleQuote: stringBuilder.Append("""); break; default: stringBuilder.Append(value[i]); break; // < and > pass through raw ``
PoC
The following program demonstrates that AngleSharp’s parser misses the injected ` element. A sanitizer walking this DOM would see nothing dangerous, yet the serialized output re-parses in a browser as a live trigger. `cs using System; using System.Linq; using AngleSharp.Html.Parser; public class Program { static readonly string Payload1 = "" + "" + "<a encoding=\"\">" + ""; public static void Main() { var parser = new HtmlParser(); Check(parser, Payload1, "IMG", "AngleSharp missed – VULNERABLE (mXSS via attribute serialization)", "AngleSharp found – SAFE"); } static void Check(HtmlParser parser, string html, string tag, string failMsg, string passMsg) { var doc = parser.ParseDocument(html); var tags = doc.All.Select(e => e.TagName).ToHashSet(); var found = tags.Contains(tag); Console.WriteLine(found ? passMsg : failMsg); Console.WriteLine("Serialized output:"); Console.WriteLine(doc.DocumentElement.OuterHtml); } } ``
Output: `` AngleSharp missed – VULNERABLE (mXSS via attribute serialization) Serialized output: <a encoding=""> ``
_The title tag may be swapped out for style and other RCDATA elements._
When a browser receives this string and parses annotation-xml encoding="text/html" as an HTML integration point, the ` closes the title element and the ` fires its onerror handler.
Impact
Implemented HTML sanitizers that depend and trust AngleSharp's ability to parse HTML correctly may be bypassable, as AngleSharp fails to acknowledge certain vectors under certain conditions.
This reduces AngleSharp's credibility as a conformant HTML parser.
Affected packages
Versions sourced from the GitHub Security Advisory.
| Package | Affected versions | Patched versions |
|---|---|---|
AngleSharpNuGet | < 1.5.0 | 1.5.0 |
Affected products
2Patches
Vulnerability mechanics
References
3News mentions
0No linked articles in our index yet.