# GlobalKTable consistency, ensuring latest state for processing

**URL:** <https://forum.confluent.io/t/globalktable-consistency-ensuring-latest-state-for-processing/10961>\
**Category:** Kafka Streams\
**Created:** [26 June 2024 14:55 UTC](https://forum.confluent.io/t/globalktable-consistency-ensuring-latest-state-for-processing/10961 "2024-06-26T14:55:03Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![windymindy](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/windymindy/32/3851_2.png) [@windymindy](https://forum.confluent.io/u/windymindy)\
**Post date:** [26 June 2024 14:55 UTC](https://forum.confluent.io/t/globalktable-consistency-ensuring-latest-state-for-processing/10961/1 "2024-06-26T14:55:03Z")

</div>

Hey.  
Can’t get my head around something.  
Say, one is to deduplicate a raw event stream from redundant sources.  
That is by stateful **flatTransformValues**.  
Without going into detail on what the duplicate record criteria is, let’s say one is to store some view or aggregate of all previous values for key in a GlobalKTable and return null from the transformer for duplicates.

Question.  
How is it ensured, that by the time a record is being processed the table actually has the latest state and won’t allow a duplicate through?  
What if there were two apllication nodes, one died right after processing a record, repartition happened, and the other node just got a record to process, but not the latest state?  
Is there some internal mechanism that rechekes linked partitions of service topics before giving control to the library user?  
Or is there only eventual consistency? Would one need a strongly consistent ACID storage like an SQL database?

---

<div class="post-metadata">

**Author:** ![mjsax](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/mjsax/32/3113_2.png) [@mjsax](https://forum.confluent.io/u/mjsax)\
**Post date:** [27 June 2024 06:24 UTC](https://forum.confluent.io/t/globalktable-consistency-ensuring-latest-state-for-processing/10961/2 "2024-06-27T06:24:09Z")

</div>

GlobalKTables are inherenty async, because they only update by reading from their input topic.

For de-duplication, you would need to use a regular store/KTable and partition the data accordingly. This way, your de-duplication step can read a input record, lookup the store, and update the store sync, avoiding any race condition.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex019/uploads/confluentcommunity/original/1X/c49438c90c9df282e9996fdf6971be890c71b65a.svg) [@system](https://forum.confluent.io/u/system)\
**Post date:** [4 July 2024 06:25 UTC](https://forum.confluent.io/t/globalktable-consistency-ensuring-latest-state-for-processing/10961/3 "2024-07-04T06:25:07Z")

</div>

This topic was automatically closed 7 days after the last reply. New replies are no longer allowed.
