[HN Gopher] How (not) to sign a JSON object (2019)
       ___________________________________________________________________
        
       How (not) to sign a JSON object (2019)
        
       Author : maple3142
       Score  : 36 points
       Date   : 2025-02-09 14:38 UTC (8 hours ago)
        
 (HTM) web link (www.latacora.com)
 (TXT) w3m dump (www.latacora.com)
        
       | DarkUranium wrote:
       | Not sure if I'm just misunderstanding the article or not, but it
       | feels like an overengineered solution, reminescent of SAML's
       | replacement instructions (just a hardcoded and admittedly _way_
       | better option --- but still in a similar vein of  "text
       | replacement hacks").
       | 
       | I know it's not the most elegant thing ever, but if it _needs_ to
       | be JSON at the post-signing level, why not just something like `[
       | "75cj8hgmRg+v8AQq3OvTDaf8pEWEOelNHP2x99yiu3Y","{\"foo\":\"bar\"}"
       | ]`, in other words, encode the JSON being signed as a string.
       | This would then ensure that, even if the "outer" JSON is parsed
       | and re-encoded, the string is unmodified. It'll even survive
       | weird parsing and re-encoding, which the regex replacement option
       | might not (unless it's tolerant of whitespace changes).
       | 
       | (or, for the extra paranoid: encode the latter to base64 first
       | and _then_ as a string, yielding something like `[ "75cj8hgmRg+v8
       | AQq3OvTDaf8pEWEOelNHP2x99yiu3Y","eyJmb28iOiJiYXIifQ"]` --- this
       | way, it doesn't look like JSON anymore, for any parsers that try
       | to be too smart)
       | 
       | If the outer needs to be an object (as opposed to array), this is
       | also trivially adapted, of course: `{"hmac":"75cj8hgmRg+v8AQq3OvT
       | Daf8pEWEOelNHP2x99yiu3Y","json":"{\"foo\":\"bar\"}"}`.
        
         | 38 wrote:
         | Json encoded as a string is cursed, no one should do that and
         | stop suggesting it. Base64 is fine or even ascii85
        
           | Groxx wrote:
           | base64 is often _even larger_ than an escaped JSON string.
           | and not human-readable at all.
           | 
           | I'll take stringified json-in-json 90% of the time, thanks.
           | if you're using JSON you're already choosing an inefficient,
           | human-oriented language anyway, a small bit more overhead
           | doesn't hurt.
           | 
           | (obviously neither of these are good options, just defer your
           | parsing so you retain the exact byte sequences while
           | checking, and then parse the substring. you shouldn't be
           | parsing before checking anyway. but when you can't trust
           | people to do that...)
        
           | benatkin wrote:
           | The comment you replied to was posted in good faith AFAICT.
           | Your "stop suggesting it" is unnecessarily antagonistic.
        
         | mbreese wrote:
         | There are many ways to represent the JSON as binary... and all
         | are equally valid. The easiest case to think about is with and
         | without whitespace. Because what HMAC cares about are the
         | byte[] values, not alphanumeric tokens.
         | 
         | Then, if you couple this with sending data through a proxy
         | (maybe invisible to the developers), which may or may not alter
         | that text representation, you end up with a mess. If you base64
         | encode the JSON, you now lose any benefit you might gain from
         | those intermediate proxies, as they can't read the payload...
        
         | Someone wrote:
         | > in other words, encode the JSON being signed as a string.
         | This would then ensure that, even if the "outer" JSON is parsed
         | and re-encoded, the string is unmodified. It'll even survive
         | weird parsing and re-encoding, which the regex replacement
         | option might not (unless it's tolerant of whitespace changes).
         | 
         | Would it be guaranteed to survive even standard parsing?
         | 
         | It wouldn't surprise me at all, for example, if there are json
         | parsers out there that, on reading, map "\u0009" and "\t" to
         | the same string, so that they can only round-trip one of those
         | strings. Similarly, there's the pair of "\uabcd" and "\uABCD".
         | There probably are others.
        
           | LegionMammal978 wrote:
           | Presumably when receiving the object, you'd first unescape
           | the string (which should yield a unique output unless you
           | have big parser bugs), check the UTF-8 bytes of the unescaped
           | string against the signature, and only then decode the
           | unescaped string as the inner JSON object. It shouldn't
           | matter how exactly the string is escaped, as long as it can
           | be unescaped successfully.
        
         | spankalee wrote:
         | Thats no different than the suggestion at the beginning of the
         | article to serialize the JSON and sign the string.
        
         | theamk wrote:
         | You can and this will be simple and reliable.. but that's
         | solving the different (and easier) problem that the post. In
         | the post, author wants to have still have parsable JSON _and_ a
         | signature. Think middleware which can check signature, but
         | cannot alter the contents, followed by backend expecting nice
         | JSON. Or a logging middleware which looks at individual fields.
         | Or a load balancer which checks the "user" and "project"
         | fields. Or a WAF checking for right fields. In other words:
         | 
         | > Anyone who cares about validating the signature can, and
         | anyone who cares that the JSON object has a particular
         | structure doesn't break (because the blob is still JSON and it
         | still has the data it's supposed to have in all the familiar
         | places).
         | 
         | As author mentions, you can compromise by having "hmac", "json"
         | and "user" (for routing purposes only), but this will increase
         | overall size. This is approach 2 in the blog.
        
       | treyd wrote:
       | This is one of the deeply dissatisfying parts of the Matrix spec.
       | They didn't have any constraints forcing them to embed signatures
       | within the json objects, but they elected to invent their own
       | signing scheme and do it anyways, despite being a greenfield. It
       | also includes support for a special "unsigned" portion for extra
       | data that comes along (which is often used for the server to
       | inject the age of an event).
       | 
       | I don't _think_ the protocol still injects the signature into
       | event structures, but this weird  "unsigned" field is still there
       | looking at the source json for a message I sent today, but it's
       | possible it's removed after processing and Fluffychat is just
       | removing it.
        
       | Zamicol wrote:
       | We addressed these concerns while developing Coze, a
       | cryptographic JSON messaging specification. The specification
       | details how we chose to address these concerns.
       | 
       | https://github.com/Cyphrme/Coze
        
         | er4hn wrote:
         | I saw from the main page that you are aware of COSE (RFC 8152),
         | with it's super similar name, but I didn't see anything in
         | https://github.com/Cyphrme/CozeX/blob/master/coze_vs.md
         | comparing it or CBOR.
         | 
         | Is the improvement COZE has over COSE that the body is default
         | human readable, whereas COSE it's in some machine format that
         | needs a reader util?
        
           | Zamicol wrote:
           | We picked the name Coze as a play on words on JOSE ("Cypher
           | JOSE", when Jose is typically pronounced in Spanish, "ho-
           | zay". The English word coze meaning a friendly chat was too
           | perfect for the name of a messaging specification). I
           | somewhat regret not using the three letter "coz".
           | 
           | Cose also picked its name while thinking of JOSE. Cose is
           | binary oriented, and attempts to be as similar to JOSE as
           | possible.
           | 
           | Coze is a first principles reimagining of signing JSON.
           | 
           | As an exercise, we've played around with creating a binary
           | format, (which we're calling Booze, Binary Oriented cOZE) but
           | since Coze is already much more space efficient, it's not as
           | beneficial. It may become more relevant with post quantum, as
           | currently post quantum systems are much larger than ECDSA.
           | The only other advantage would be removing the JSON
           | semantics, but that's at the cost of implementing a binary
           | format. The human readability aspect is paramount for our
           | application, and we feel it's generally better practice.
        
             | er4hn wrote:
             | Thanks for explaining the background of that.
             | 
             | I agree that human readability can make things somewhat
             | easier. The part about a binary format for the signatures
             | confuses me a little - the signatures are going to be big
             | no matter what. If you care about minimizing size would it
             | make sense to just find a more efficient way to encode the
             | signatures rather than change the entire message format?
        
               | Zamicol wrote:
               | Coze uses base64 encoding for binary values such as `tmb`
               | and `sig`. `tmb` isn't a problem since digests are
               | designed to be short, but signatures for some primitive
               | might be very large.
               | 
               | When compared to encoding a value directly in binary,
               | base64 has about a 25% overhead (6 bits /8 bits, 3/4). As
               | far as the concern about using better encoding, base64 is
               | just about as good as it gets while being maximally
               | compatible. If using base 128 (7 bit ASCII), there's too
               | many incompatible special characters for a human readable
               | format. The full 8/8 bits, extended ASCII, isn't
               | generally possible as systems use UTF-8 which begins
               | using multiple bytes. (I've done a lot of work in this
               | area, including a patent on base conversion. See also
               | convert.zamicol.com) An advantage of a binary format is
               | that there is minimal encoding overhead for binary values
               | (escaping/padding is typically the only overhead, so
               | usually around 99% efficient compared to base64's 75%.)
               | 
               | This isn't too much of a concern when signatures are
               | small as encoding inefficiency is small compared to the
               | payload's overall size, but if signatures are in the
               | kilobytes or even megabytes, that extra 25% becomes
               | meaningful for some hyper-efficient applications, like
               | high cost blockchains. Our thought is using post quantum
               | is already much more massive than existing elliptic
               | curve, so any future applications of post quantum are
               | going to have to deal with much larger signatures
               | anyways. The signatures can also be stored on disk using
               | binary or compressed which also makes it not a concern.
        
               | tzot wrote:
               | > When compared to encoding a value directly in binary,
               | base64 has about a 25% overhead (6 bits /8 bits, 3/4).
               | 
               | The original data take >=25% less space than the
               | base64-encoded data, but the base64-encoded data take
               | >=33 1/3 % more space than the original data. The
               | overhead is about 33 1/3 %.
        
         | j-krieger wrote:
         | Very cool! It shares its name with "COSE" (RFC 8152), a signing
         | scheme for CBOR objects :)
        
       | codeflo wrote:
       | I have a feeling that idea 2 is a recipe for disaster:
       | 
       | > Add the tag and the exact string you signed to the object,
       | validate the signature and then validate that the JSON object is
       | the same as the one you got.
       | 
       | In cryptographic practice, redundant information usually spells
       | disaster, because inevitably, someone will use the copy that
       | wasn't verified.
       | 
       | But let's dig into it. If I understand this correctly, the
       | suggestion is to have something like this:                   {
       | "object": "with",             "some": "properties",
       | "signatureInfo": {                 "signedString":
       | "{\"object\":\"with\",\"some\":\"properties\"}",
       | "signature": "... base64-encoded-signature ..."             }
       | }
       | 
       | It's mentioned that "the downside is your messages are about
       | twice the size that they need to be". In my opinion, this scheme
       | is pointless. To verify "the JSON object is the same as the one
       | you got", you have to do what?
       | 
       | 1. Parse the outer object as JSON, extract and remove
       | signatureInfo.
       | 
       | 2. Verify the signature.
       | 
       | 3. Parse the signedString as JSON.
       | 
       | 4. Verify that the object you got in step 1 equal to object you
       | got in step 3 using a some kind of deep equality.
       | 
       | First of all, this is error prone, and as underspecified as JSON
       | is, there are potential exploits if the comparison isn't done
       | carefully. But even worse, if you think about it, the outer JSON
       | is entirely useless, since you need to parse the inner JSON
       | anyway -- so why not just use it directly?
       | 
       | It seems to me that this suggestion is strictly worse than just
       | sending the inner part:                   {
       | "signedString": "{\"object\":\"with\",\"some\":\"properties\"}",
       | "signature": "... base64-encoded-signature ..."         }
       | 
       | Yes, it's no longer "in-band", but I don't think it was really
       | in-band before, it was just out-of-band with an outer layer of
       | redundant information.
        
       | saurik wrote:
       | Supposedly, this is the non-code documentation for AWS Version 3
       | signing.
       | 
       | https://docs.aws.amazon.com/amazonswf/latest/developerguide/...
        
       | askvictor wrote:
       | Another problem with signing JSON: you can have two different
       | json objects that mean the same thing, and will do exactly the
       | same thing in your code e.g. {"a": "foo", "b": "bar"} vs {"b":
       | "bar", "a": "foo"}. Also, whitespace. Are there any standards for
       | normalising json, so that two equivalent, but differently written
       | JSON files will have the same signature?
        
         | busymom0 wrote:
         | This is why, in one of my projects, I first stringified the
         | JSON using built in JSON.stringify(your_json) function, then
         | signed that string and sent the string, its signature, and
         | public key to server. Server verifies the signature using the
         | string and if passes, then uses JSON.parse(your_string) to get
         | the original json.
        
         | lxgr wrote:
         | That's canonicalization, and the article does mention it (but
         | unfortunately does not offer much insight other than that it's
         | hard).
        
       | lxgr wrote:
       | > Unless you have a good reason why you need an (asymmetric)
       | signature, you want a MAC.
       | 
       | Is "I want the server/validating side to be safe even against
       | server-side attackers with read-only permissions" not a good
       | reason? Because that's one thing that asymmetric signatures
       | provide out of the box compared to MACs.
        
       | Muromec wrote:
       | Or just use asn1 like _normal_ people.
        
       ___________________________________________________________________
       (page generated 2025-02-09 23:01 UTC)