[HN Gopher] How (not) to sign a JSON object (2019)
___________________________________________________________________
How (not) to sign a JSON object (2019)
Author : maple3142
Score : 36 points
Date : 2025-02-09 14:38 UTC (8 hours ago)
(HTM) web link (www.latacora.com)
(TXT) w3m dump (www.latacora.com)
| DarkUranium wrote:
| Not sure if I'm just misunderstanding the article or not, but it
| feels like an overengineered solution, reminescent of SAML's
| replacement instructions (just a hardcoded and admittedly _way_
| better option --- but still in a similar vein of "text
| replacement hacks").
|
| I know it's not the most elegant thing ever, but if it _needs_ to
| be JSON at the post-signing level, why not just something like `[
| "75cj8hgmRg+v8AQq3OvTDaf8pEWEOelNHP2x99yiu3Y","{\"foo\":\"bar\"}"
| ]`, in other words, encode the JSON being signed as a string.
| This would then ensure that, even if the "outer" JSON is parsed
| and re-encoded, the string is unmodified. It'll even survive
| weird parsing and re-encoding, which the regex replacement option
| might not (unless it's tolerant of whitespace changes).
|
| (or, for the extra paranoid: encode the latter to base64 first
| and _then_ as a string, yielding something like `[ "75cj8hgmRg+v8
| AQq3OvTDaf8pEWEOelNHP2x99yiu3Y","eyJmb28iOiJiYXIifQ"]` --- this
| way, it doesn't look like JSON anymore, for any parsers that try
| to be too smart)
|
| If the outer needs to be an object (as opposed to array), this is
| also trivially adapted, of course: `{"hmac":"75cj8hgmRg+v8AQq3OvT
| Daf8pEWEOelNHP2x99yiu3Y","json":"{\"foo\":\"bar\"}"}`.
| 38 wrote:
| Json encoded as a string is cursed, no one should do that and
| stop suggesting it. Base64 is fine or even ascii85
| Groxx wrote:
| base64 is often _even larger_ than an escaped JSON string.
| and not human-readable at all.
|
| I'll take stringified json-in-json 90% of the time, thanks.
| if you're using JSON you're already choosing an inefficient,
| human-oriented language anyway, a small bit more overhead
| doesn't hurt.
|
| (obviously neither of these are good options, just defer your
| parsing so you retain the exact byte sequences while
| checking, and then parse the substring. you shouldn't be
| parsing before checking anyway. but when you can't trust
| people to do that...)
| benatkin wrote:
| The comment you replied to was posted in good faith AFAICT.
| Your "stop suggesting it" is unnecessarily antagonistic.
| mbreese wrote:
| There are many ways to represent the JSON as binary... and all
| are equally valid. The easiest case to think about is with and
| without whitespace. Because what HMAC cares about are the
| byte[] values, not alphanumeric tokens.
|
| Then, if you couple this with sending data through a proxy
| (maybe invisible to the developers), which may or may not alter
| that text representation, you end up with a mess. If you base64
| encode the JSON, you now lose any benefit you might gain from
| those intermediate proxies, as they can't read the payload...
| Someone wrote:
| > in other words, encode the JSON being signed as a string.
| This would then ensure that, even if the "outer" JSON is parsed
| and re-encoded, the string is unmodified. It'll even survive
| weird parsing and re-encoding, which the regex replacement
| option might not (unless it's tolerant of whitespace changes).
|
| Would it be guaranteed to survive even standard parsing?
|
| It wouldn't surprise me at all, for example, if there are json
| parsers out there that, on reading, map "\u0009" and "\t" to
| the same string, so that they can only round-trip one of those
| strings. Similarly, there's the pair of "\uabcd" and "\uABCD".
| There probably are others.
| LegionMammal978 wrote:
| Presumably when receiving the object, you'd first unescape
| the string (which should yield a unique output unless you
| have big parser bugs), check the UTF-8 bytes of the unescaped
| string against the signature, and only then decode the
| unescaped string as the inner JSON object. It shouldn't
| matter how exactly the string is escaped, as long as it can
| be unescaped successfully.
| spankalee wrote:
| Thats no different than the suggestion at the beginning of the
| article to serialize the JSON and sign the string.
| theamk wrote:
| You can and this will be simple and reliable.. but that's
| solving the different (and easier) problem that the post. In
| the post, author wants to have still have parsable JSON _and_ a
| signature. Think middleware which can check signature, but
| cannot alter the contents, followed by backend expecting nice
| JSON. Or a logging middleware which looks at individual fields.
| Or a load balancer which checks the "user" and "project"
| fields. Or a WAF checking for right fields. In other words:
|
| > Anyone who cares about validating the signature can, and
| anyone who cares that the JSON object has a particular
| structure doesn't break (because the blob is still JSON and it
| still has the data it's supposed to have in all the familiar
| places).
|
| As author mentions, you can compromise by having "hmac", "json"
| and "user" (for routing purposes only), but this will increase
| overall size. This is approach 2 in the blog.
| treyd wrote:
| This is one of the deeply dissatisfying parts of the Matrix spec.
| They didn't have any constraints forcing them to embed signatures
| within the json objects, but they elected to invent their own
| signing scheme and do it anyways, despite being a greenfield. It
| also includes support for a special "unsigned" portion for extra
| data that comes along (which is often used for the server to
| inject the age of an event).
|
| I don't _think_ the protocol still injects the signature into
| event structures, but this weird "unsigned" field is still there
| looking at the source json for a message I sent today, but it's
| possible it's removed after processing and Fluffychat is just
| removing it.
| Zamicol wrote:
| We addressed these concerns while developing Coze, a
| cryptographic JSON messaging specification. The specification
| details how we chose to address these concerns.
|
| https://github.com/Cyphrme/Coze
| er4hn wrote:
| I saw from the main page that you are aware of COSE (RFC 8152),
| with it's super similar name, but I didn't see anything in
| https://github.com/Cyphrme/CozeX/blob/master/coze_vs.md
| comparing it or CBOR.
|
| Is the improvement COZE has over COSE that the body is default
| human readable, whereas COSE it's in some machine format that
| needs a reader util?
| Zamicol wrote:
| We picked the name Coze as a play on words on JOSE ("Cypher
| JOSE", when Jose is typically pronounced in Spanish, "ho-
| zay". The English word coze meaning a friendly chat was too
| perfect for the name of a messaging specification). I
| somewhat regret not using the three letter "coz".
|
| Cose also picked its name while thinking of JOSE. Cose is
| binary oriented, and attempts to be as similar to JOSE as
| possible.
|
| Coze is a first principles reimagining of signing JSON.
|
| As an exercise, we've played around with creating a binary
| format, (which we're calling Booze, Binary Oriented cOZE) but
| since Coze is already much more space efficient, it's not as
| beneficial. It may become more relevant with post quantum, as
| currently post quantum systems are much larger than ECDSA.
| The only other advantage would be removing the JSON
| semantics, but that's at the cost of implementing a binary
| format. The human readability aspect is paramount for our
| application, and we feel it's generally better practice.
| er4hn wrote:
| Thanks for explaining the background of that.
|
| I agree that human readability can make things somewhat
| easier. The part about a binary format for the signatures
| confuses me a little - the signatures are going to be big
| no matter what. If you care about minimizing size would it
| make sense to just find a more efficient way to encode the
| signatures rather than change the entire message format?
| Zamicol wrote:
| Coze uses base64 encoding for binary values such as `tmb`
| and `sig`. `tmb` isn't a problem since digests are
| designed to be short, but signatures for some primitive
| might be very large.
|
| When compared to encoding a value directly in binary,
| base64 has about a 25% overhead (6 bits /8 bits, 3/4). As
| far as the concern about using better encoding, base64 is
| just about as good as it gets while being maximally
| compatible. If using base 128 (7 bit ASCII), there's too
| many incompatible special characters for a human readable
| format. The full 8/8 bits, extended ASCII, isn't
| generally possible as systems use UTF-8 which begins
| using multiple bytes. (I've done a lot of work in this
| area, including a patent on base conversion. See also
| convert.zamicol.com) An advantage of a binary format is
| that there is minimal encoding overhead for binary values
| (escaping/padding is typically the only overhead, so
| usually around 99% efficient compared to base64's 75%.)
|
| This isn't too much of a concern when signatures are
| small as encoding inefficiency is small compared to the
| payload's overall size, but if signatures are in the
| kilobytes or even megabytes, that extra 25% becomes
| meaningful for some hyper-efficient applications, like
| high cost blockchains. Our thought is using post quantum
| is already much more massive than existing elliptic
| curve, so any future applications of post quantum are
| going to have to deal with much larger signatures
| anyways. The signatures can also be stored on disk using
| binary or compressed which also makes it not a concern.
| tzot wrote:
| > When compared to encoding a value directly in binary,
| base64 has about a 25% overhead (6 bits /8 bits, 3/4).
|
| The original data take >=25% less space than the
| base64-encoded data, but the base64-encoded data take
| >=33 1/3 % more space than the original data. The
| overhead is about 33 1/3 %.
| j-krieger wrote:
| Very cool! It shares its name with "COSE" (RFC 8152), a signing
| scheme for CBOR objects :)
| codeflo wrote:
| I have a feeling that idea 2 is a recipe for disaster:
|
| > Add the tag and the exact string you signed to the object,
| validate the signature and then validate that the JSON object is
| the same as the one you got.
|
| In cryptographic practice, redundant information usually spells
| disaster, because inevitably, someone will use the copy that
| wasn't verified.
|
| But let's dig into it. If I understand this correctly, the
| suggestion is to have something like this: {
| "object": "with", "some": "properties",
| "signatureInfo": { "signedString":
| "{\"object\":\"with\",\"some\":\"properties\"}",
| "signature": "... base64-encoded-signature ..." }
| }
|
| It's mentioned that "the downside is your messages are about
| twice the size that they need to be". In my opinion, this scheme
| is pointless. To verify "the JSON object is the same as the one
| you got", you have to do what?
|
| 1. Parse the outer object as JSON, extract and remove
| signatureInfo.
|
| 2. Verify the signature.
|
| 3. Parse the signedString as JSON.
|
| 4. Verify that the object you got in step 1 equal to object you
| got in step 3 using a some kind of deep equality.
|
| First of all, this is error prone, and as underspecified as JSON
| is, there are potential exploits if the comparison isn't done
| carefully. But even worse, if you think about it, the outer JSON
| is entirely useless, since you need to parse the inner JSON
| anyway -- so why not just use it directly?
|
| It seems to me that this suggestion is strictly worse than just
| sending the inner part: {
| "signedString": "{\"object\":\"with\",\"some\":\"properties\"}",
| "signature": "... base64-encoded-signature ..." }
|
| Yes, it's no longer "in-band", but I don't think it was really
| in-band before, it was just out-of-band with an outer layer of
| redundant information.
| saurik wrote:
| Supposedly, this is the non-code documentation for AWS Version 3
| signing.
|
| https://docs.aws.amazon.com/amazonswf/latest/developerguide/...
| askvictor wrote:
| Another problem with signing JSON: you can have two different
| json objects that mean the same thing, and will do exactly the
| same thing in your code e.g. {"a": "foo", "b": "bar"} vs {"b":
| "bar", "a": "foo"}. Also, whitespace. Are there any standards for
| normalising json, so that two equivalent, but differently written
| JSON files will have the same signature?
| busymom0 wrote:
| This is why, in one of my projects, I first stringified the
| JSON using built in JSON.stringify(your_json) function, then
| signed that string and sent the string, its signature, and
| public key to server. Server verifies the signature using the
| string and if passes, then uses JSON.parse(your_string) to get
| the original json.
| lxgr wrote:
| That's canonicalization, and the article does mention it (but
| unfortunately does not offer much insight other than that it's
| hard).
| lxgr wrote:
| > Unless you have a good reason why you need an (asymmetric)
| signature, you want a MAC.
|
| Is "I want the server/validating side to be safe even against
| server-side attackers with read-only permissions" not a good
| reason? Because that's one thing that asymmetric signatures
| provide out of the box compared to MACs.
| Muromec wrote:
| Or just use asn1 like _normal_ people.
___________________________________________________________________
(page generated 2025-02-09 23:01 UTC)