# A Small Schema Mistake That Slowed Down My Chat System (And How I Fixed It)

While developing my chat backend, I initially made a design decision that seemed ideal untill it didn't. As the system scaled, it led to significant performance issues.

It wasn’t a complex bug.  
It was a **schema design mistake**.

Let's explore the mistake, how I resolved it, and the trade-offs involved.

# The Feature: Read Receipts

![](https://cdn.hashnode.com/uploads/covers/64f2e4b017b66c6d6d7201fe/5c4004b2-9ef0-4d8e-9d63-fbf9ab980472.png align="center")

Read receipts are simple in concept but very visible in practice:

*   In 1:1 chats → you see if the other person read your message.
    
*   In group chats → you might want to know which participants have seen a message.
    

I wanted a straightforward way to answer “who has seen this message?” on a per-message basis.

## My Initial Approach

I went with a simple and intuitive design.😀

Each message stored a `readBy` array:

```json
{ 
    "chatId": "chat_1", 
    "senderId": "user_1", 
    "content": "hello", 
    "messageType": "text", 
    "readBy": [ 
      { 
       "userId": "user_1", 
       "readAt": "2025-11-24T15:54:01Z" 
      }, 
      { 
        "userId": "user_2", 
        "readAt": "2025-11-24T15:54:01Z" 
      } 
    ], 
    "createdAt": "2025-11-24T15:54:01Z" 
}
```

Whenever a user read the message, their data was pushed into this array.

This felt right: each message contains its own read state. No joins, one document to fetch the message and who read it.

### Why This Worked Initially

In 1:1 chats, it was perfect. Only two users ever wrote to that message document, the array stayed tiny, and queries were fast.

Things broke when I introduced group chats.

Imagine a group with 100 members.  
Every time a message is read, 100 users update the same document.

## The Problems

### 1\. Hot Document Problem

All users were updating the **same message document** causing

*   Write contention
    
*   Slower API responses
    

### 2\. Document Size Growth

The `readBy` array kept increasing:

*   More users → larger array
    
*   More memory + disk usage
    
*   Risk of hitting MongoDB document limits (16MB)
    

### 3\. Performance Degradation

*   Every read causes a write to the message doc, even though that write is tiny semantically.
    
*   Indexes on large arrays become heavier.
    
*   Returning “who read this message” for many messages becomes expensive.
    

In short: a design that was fine for two participants suddenly produced lots of small concurrent writes and large shared documents when groups grew.😑

## Realization

The problem wasn’t my API logic.  
It was **where I was storing the data**.

I was coupling:

> message data + user-specific read state

Which doesn’t scale well.

I had to rethink my schema design choices as I was coupling too much data which should have been separate from the start.

(One more instance where I was coupling was storing the members array in the chats document).

## The Fix: Rethinking the Schema

Instead of storing read data inside messages, I created a **separate collection**.

I decoupled the data of users and chats. With this new **Membership** collection, i was storing the data related to a user for a specific chat.

Let us see examples of the documents in action.

We are concerned with mainly 4 documents here: **User, Chat, Message** and **Membership.**

**User**

```json
{ 
    "userId": "user_1", 
    "name": "Mohit Pandey", 
    "email": "pandeymohit215@gmail.com", 
    "createdAt": "2025-11-24T15:54:01Z" 
}
```

**Chat**

```json
{ 
    "chatId": "chat_1", 
    "type": "group",
    "name": "System Design Cohort", 
    "messageSeqCount":562, 
    "createdAt": "2025-11-24T15:54:01Z" 
}
```

**Message**

```json
{ 
    "chatId": "chat_1", 
    "senderId": "user_1", 
    "content": "hello", 
    "messageType": "text", 
    "messageSeq":566,
    "createdAt": "2025-11-24T15:54:01Z" 
}
```

**Membership**

```json
{ 
    "chatId": "chat_1", 
    "user_id": "user_1",
    "lastReadMessageSeq":545, 
    "createdAt": "2025-11-24T15:54:01Z" 
}
```

## How This Works

By storing the `messageSeqCount` , `lastReadMessageSeq` and `messageSeq` , to get the unread messages, we simply do

`unreadCount = messageSeqCount - lastReadMessageSeq`

Whenever a new message is sent, the `messageSeqCount` is modified. This ensures the tracking of unread messages.

\*Note: This method is not suitable if we want the "\****Hard Delete"*** *feature as it would alter the message sequence. We get all the messages count even if deleted. To handle this, we flag the message as deleted and modify the UI based on this flag.*

## Why This Works Better

### No Large Arrays

No more `readBy` or members stored inside a single document.  
This avoids unbounded growth and eliminates frequent updates to the same record.

### No Write Contention

Each user updates their own membership record instead of a shared document.  
This removes the “hot document” problem entirely.

### Faster APIs

No repeated reads and writes on the same document → less locking, quicker responses.

### Scales with Group Size

100 users = 100 independent updates, not a single bottleneck.  
The system scales naturally as more users are added.

## Trade-offs

This approach isn’t perfect.

*   It requires an extra collection, hence, for a large group chat, the number of documents also increase.
    
*   Fetching chat data may require additional queries or joins to combine messages and membership data.
    
*   Maintaining `messageSeq` requires careful handling to ensure consistency and atomicity during writes.
    

But the benefits far outweigh the costs. In most real-world systems, avoiding write contention is far more critical than reducing the number of queries.

## Mitigating The Trade-offs

1.  **Redis Caching:** We cache the user’s inbox (chat list) in Redis to avoid running expensive joins or `$lookup` queries on every request. This significantly improves read performance.
    
2.  Denormalising the `memberCount` and storing it in the chat document can help us to just show the badge instead of showing all members if we do not have to.
    
3.  We can add database indexing to enable faster access of data such as compound index for userId and chatId in Membership table.
    

## Key Learning

This crucial shift in my design helped me build stable and scalable apis. This made me understand that a design that works at small scale can fail badly at large scale.

Especially when:

*   Multiple users write to the same document
    
*   Data keeps growing inside a single record
    

## Conclusion

What looked like a small schema change ended up making a significant difference in performance and scalability.

This experience made one thing very clear — building systems isn’t just about making things work, it’s about making them work **efficiently under scale**.

When dealing with real-time systems or multiple users, it’s important to think beyond the immediate implementation and consider:

*   **Write patterns** → who is updating what, and how often
    
*   **Data growth** → how your data structures evolve over time
    
*   **Contention points** → where multiple operations might compete
    

The biggest takeaway for me was simple:

> Good performance doesn’t come from optimizing queries later — it comes from making the right data modeling decisions early.🙌

Do share your views about this choice and how you would have implemented this. Thank you.🤝
