# Does the same usda file translate into the same usdc data?

**URL:** <https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499>\
**Category:** USD\
**Created:** [June 4, 2025, 4:29pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499 "2025-06-04T16:29:40Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fanny](https://avatars.discourse-cdn.com/v4/letter/f/b38774/32.png) [@Fanny](https://forum.aousd.org/u/Fanny)\
**Post date:** [June 4, 2025, 4:29pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/1 "2025-06-04T16:29:40Z")

</div>

Hello everyone, I’m trying to find out if a layer has changed by looking at its usdc data, but I’m finding that for the same content (the same usda) I can get two different usc data.  
I’ve put up this little script that creates two layers that have the same prim. If I clear the second layer and recreate the same prim and save it I get two different hashes although the content of the two layers is the same.

```auto
import hashlib
from pxr import Sdf

filepath1 = 'file1.usdc'
layer1 = Sdf.Layer.CreateNew(filepath1)
Sdf.CreatePrimInLayer(layer1, '/world')
layer1.Save()

with open(filepath1, 'rb') as f:
    data = f.read()
print(layer1.ExportToString())
print(data)
print(hashlib.md5(data).digest())

filepath2 = 'file2.usdc'
layer2 = Sdf.Layer.CreateNew(filepath2)
Sdf.CreatePrimInLayer(layer2, '/world')
layer2.Save()
layer2.Clear()
Sdf.CreatePrimInLayer(layer2, '/world')
layer2.Save()

with open(filepath2, 'rb') as f:
    data2 = f.read()
print(layer2.ExportToString())
print(data2)
print(hashlib.md5(data2).digest())

```

If I don’t Clear() and CreatePrimInLayer() and Save() again the hashes are identical.  
Looking at the documentation the Clear() method is undoable, so does it store extra data or timestamp info ?`https://openusd.org/release/api/class_sdf_layer.html#a9013e716d1676f98b48ab913031e6d01`

Thanks for the help

---

<div class="post-metadata">

**Author:** ![dhruvgovil](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.aousd.org/dhruvgovil/32/3_2.png) [@dhruvgovil](https://forum.aousd.org/u/dhruvgovil)\
**Post date:** [June 5, 2025, 2:24pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/2 "2025-06-05T14:24:04Z")

</div>

Crate files don’t necessarily store their data in the exact same order every time. I think you can get lucky and they might match up, but as far as I know, there’s no such guarantee in the crate writer.

As for your specific example, I think it’s small enough that it should be identical. I’ll look into it because I’m curious, but again, I don’t think you should expect repeatable file layout.

---

<div class="post-metadata">

**Author:** ![dhruvgovil](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.aousd.org/dhruvgovil/32/3_2.png) [@dhruvgovil](https://forum.aousd.org/u/dhruvgovil)\
**Post date:** [June 5, 2025, 4:12pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/3 "2025-06-05T16:12:32Z")

</div>

So I like using [ImHex](https://github.com/WerWolv/ImHex) to investigate and diff files, and it gives me the following diff.

 ![image](https://us1.discourse-cdn.com/flex016/uploads/aousd/original/1X/1fe97994e82acab2a218268131b1a8933aec88cf.png)

To my eye that is showing that there’s a difference in where some padding is going between writes, which is changing where the offsets point to.

Otherwise when parsed, this is the dictionary output of the two files, which is identical.

**file1.usdc**

```json
{'/': (PseudoRoot, {'primChildren': TokenVector:['world']}),
 '/world': (Prim, {})}

```

**file2.usdc**

```json
{'/': (PseudoRoot, {'primChildren': TokenVector:['world']}),
 '/world': (Prim, {})}

```

@alexmohr might have some more thoughts here, but I think its just down to the order shifts between writes.

---

<div class="post-metadata">

**Author:** ![amohr](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@amohr](https://forum.aousd.org/u/amohr)\
**Post date:** [June 5, 2025, 5:20pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/4 "2025-06-05T17:20:11Z")

</div>

Right – in addition to padding bits we use hash tables in the .usdc implementation, and don’t take pains (or the perf cost) to sort everything when we write, so just doing `usdcat` on a `.usdc` file is liable to produce results that differ bitwise, but not content-wise.

Even if we did ensure that `usdcat` with a `.usdc` always produced bitwise-identical results, you can still run into trouble because an incremental `Save()` of a `.usdc` does not in general rewrite the whole file. It’s sort of “journaled”. So an edited and `Save()`d `.usdc` file would still differ bitwise from that same file run through `usdcat`.

I think if we want content fingerprint hashing, it would be best to do it at the `SdfLayer` level, so that it would work automatically and identically for every file format that USD understands, and would avoid any issues with `.usdc` or any other “database-esque” formats. And you could do things like freely flip/flop your `.usd` file between text and binary without changing the content fingerprint hash.

Of course that’s a small project. Today I think the best way to get consistent results is to `usdcat` the `.usd` file to `.usda` and hash the output.

---

<div class="post-metadata">

**Author:** ![dhruvgovil](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.aousd.org/dhruvgovil/32/3_2.png) [@dhruvgovil](https://forum.aousd.org/u/dhruvgovil)\
**Post date:** [June 5, 2025, 5:23pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/5 "2025-06-05T17:23:49Z")

</div>

Thanks for the additional info, Alex. Yeah I forgot about the update write in place.

I think usddiff might help here but I haven’t looked at its implementation.

---

<div class="post-metadata">

**Author:** ![amohr](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@amohr](https://forum.aousd.org/u/amohr)\
**Post date:** [June 5, 2025, 5:31pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/6 "2025-06-05T17:31:09Z")

</div>

In short, `usddiff` just `usdcat`s to `.usda` and runs your diff tool of choice.

---

<div class="post-metadata">

**Author:** ![dhruvgovil](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.aousd.org/dhruvgovil/32/3_2.png) [@dhruvgovil](https://forum.aousd.org/u/dhruvgovil)\
**Post date:** [June 5, 2025, 5:32pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/7 "2025-06-05T17:32:33Z")

</div>

Ah that’s useful but different than I assumed.

Might be a cool future project to have an sdf diff and merge tool. Could be handy for merge conflicts etc in git for crate files.

---

<div class="post-metadata">

**Author:** ![amohr](https://avatars.discourse-cdn.com/v4/letter/a/f475e1/32.png) [@amohr](https://forum.aousd.org/u/amohr)\
**Post date:** [June 5, 2025, 5:42pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/8 "2025-06-05T17:42:05Z")

</div>

Definitely – `usddiff` plus `usdedit` is almost there – `usdedit` converts your layer to `.usda`, pops you into your editor of choice, and when you’re done writes it back in the original format, so you don’t have to think about what type of file you’re dealing with.

A USD-specific merge tool would be interesting, since it could operate not only at the bare textual level, but also at the higher structural content level. Something like a user-guided selective “flattening” where the prim hierarchy and listops and dictionaries and so on are understood.

---

<div class="post-metadata">

**Author:** ![dhruvgovil](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.aousd.org/dhruvgovil/32/3_2.png) [@dhruvgovil](https://forum.aousd.org/u/dhruvgovil)\
**Post date:** [June 6, 2025, 3:54pm UTC](https://forum.aousd.org/t/does-the-same-usda-file-translate-into-the-same-usdc-data/2499/9 "2025-06-06T15:54:47Z")

</div>

Yeah exactly. I think being able to do it at a structural level would mean that you’d not be dependent on ordering of prims for textual diffs, and also be able to handle crate files.

We’ve done similar for other structural data types in our pipelines and its super handy.
