[{"content":" Welcome to my page # ","date":"29 April 2026","externalUrl":null,"permalink":"/","section":"","summary":"","title":"","type":"page"},{"content":"Authentication becomes confusing when different topics are mixed together without a clear path. Some methods are about sending a secret. Some are about remembering a logged-in user. Some are about granting an app permission. Some are about identifying the user. And some are about making one login work across many apps.\nThe easiest way to understand it is to learn in this order:\nBasic Authentication Digest Authentication API Keys Session-based Authentication Bearer Tokens and JWT Access Tokens and Refresh Tokens OAuth 2.0 OpenID Connect Single Sign-On That order matters. Each topic is basically trying to fix a limitation of the previous one. # 1) Basic Authentication # Basic Authentication is the simplest possible idea:\nThe client sends a username and password with the request.\nUsually, the username and password are joined like this:\nusername:password Then that value is Base64-encoded and sent in the HTTP header.\nExample shape:\nAuthorization: Basic dXNlcjpwYXNz The important thing is this:\nBase64 is not security. It is just a different way to represent the same text.\nSo Basic Auth is basically:\n“Here is my username and password on this request.”\nWhy it exists # Because it is simple.\nIt works well for:\nquick testing internal tools scripts very simple systems It is easy to add and easy to understand.\nFlow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant Client participant Server Client-\u003e\u003eServer: Request + username/password (Base64 encoded) Server-\u003e\u003eServer: Decode and verify alt Valid Server--\u003e\u003eClient: Allow access else Invalid Server--\u003e\u003eClient: Reject request end Example shape # GET /profile HTTP/1.1 Host: example.com Authorization: Basic YWxpY2U6c2VjcmV0MTIz Main objects / terminologies # Username The identity claim from the client.\nPassword The secret used to prove the user knows the account secret.\nAuthorization header The HTTP header where the credentials are sent.\nBase64 encoding A text encoding format. Not encryption.\nCaveats # Password is effectively being sent with requests. Base64 does not protect the password. Should only be used over HTTPS. Not great for modern apps. No built-in concept of login session, token expiry, or logout. 2) Digest Authentication # Digest Authentication tries to improve Basic Auth.\nInstead of sending the actual password, the client sends a calculated response based on:\nthe username the password a server-provided challenge request details So the client is proving:\n“I know the password”\nwithout sending the raw password directly.\nThat is the key idea.\nWhy it exists # Basic Auth is too direct. Digest was created to avoid sending the password itself.\nSo Digest introduces a challenge-response pattern:\nserver sends a challenge client uses the password to compute a digest server verifies that digest Flow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant Client participant Server Client-\u003e\u003eServer: Request protected resource Server--\u003e\u003eClient: Challenge (nonce, realm, etc.) Client-\u003e\u003eClient: Compute digest using password + challenge Client-\u003e\u003eServer: Send digest response Server-\u003e\u003eServer: Verify digest alt Valid Server--\u003e\u003eClient: Allow access else Invalid Server--\u003e\u003eClient: Reject request end # Example shape # Authorization: Digest username=\u0026#34;alice\u0026#34;, realm=\u0026#34;admin\u0026#34;, nonce=\u0026#34;abc123\u0026#34;, uri=\u0026#34;/profile\u0026#34;, response=\u0026#34;computed-value\u0026#34; Main objects / terminologies # Challenge Data sent by the server to the client before authentication.\nNonce A server-generated value used in the calculation to reduce replay risk.\nDigest A computed hash-like response derived from the password and challenge.\nRealm A label for the protected area.\nCaveats # More secure than Basic in concept, but more complex. Still part of older HTTP auth styles. Rare in modern app architecture. Does not solve sessions, SSO, delegated access, or identity federation. Often learned for understanding history, not because it is the main modern choice. 3) API Keys # An API key is a secret string given to a client application, which the client sends to the server to identify itself.\nThis usually means:\n“This request is coming from this app/project/integration.”\nNot necessarily:\n“This request is coming from this human user.”\nThat distinction is very important.\nAPI keys are often more about client identification than real end-user authentication.\nWhy it exists # APIs need a simple way to know:\nwhich app is calling who should be rate limited which project should be billed which integration should be allowed or blocked API keys are a lightweight solution for that.\nFlow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant ClientApp participant APIServer participant PolicyStore ClientApp-\u003e\u003eAPIServer: Request + API key APIServer-\u003e\u003ePolicyStore: Check key PolicyStore--\u003e\u003eAPIServer: Valid / invalid / quota info alt Allowed APIServer--\u003e\u003eClientApp: Response else Denied APIServer--\u003e\u003eClientApp: Reject request end Example shape # GET /weather HTTP/1.1 Host: api.example.com X-API-Key: abcdef12345 Sometimes also:\nGET /weather?api_key=abcdef12345 Main objects / terminologies # API key A shared secret used by a client app to identify itself.\nClient application The app or system making the request.\nQuota / rate limit Rules controlling how much the client can use the API.\nHeader vs query parameter Common places where the key is sent.\nCaveats # Usually identifies the app, not the user. If stolen, it can often be reused until revoked. Should not be treated like a full user login system. Avoid putting it in URLs when possible. Best for simple API access control, not full identity. 4) Session-based Authentication # Session-based authentication is the classic web-app login model.\nThe user logs in once. The server then creates a session and remembers that user.\nAfter that, the browser sends a session ID on future requests, and the server uses that session ID to look up the logged-in user.\nSo the real idea is:\nthe server remembers you\nThis is why it is called stateful authentication.\nWhy it exists # Without sessions, the user would need to send credentials again and again.\nSessions solve that problem by separating:\nthe login step from later authenticated requests That makes web apps feel normal:\nlog in once click around many pages stay signed in Flow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant User participant Browser participant Server participant SessionStore User-\u003e\u003eBrowser: Enter username/password Browser-\u003e\u003eServer: Login request Server-\u003e\u003eServer: Verify credentials Server-\u003e\u003eSessionStore: Create session SessionStore--\u003e\u003eServer: session_id Server--\u003e\u003eBrowser: Set session cookie Browser-\u003e\u003eServer: Later request + session cookie Server-\u003e\u003eSessionStore: Look up session SessionStore--\u003e\u003eServer: Logged in user Server--\u003e\u003eBrowser: Protected response Example shape # Login response:\nSet-Cookie: session_id=abc123; HttpOnly; Secure; SameSite=Lax Later request:\nCookie: session_id=abc123 Main objects / terminologies # Session Server-side stored login state.\nSession ID The identifier sent by the browser so the server can find the session.\nCookie The browser mechanism commonly used to send the session ID.\nStateful The server stores and remembers session state.\nCaveats # The browser usually stores only the session ID, not the whole session. Session auth is not the same as browser sessionStorage. If the session cookie is stolen, the session may be hijacked. Cookie security matters: HttpOnly, Secure, SameSite. Great for classic web apps, but less natural for distributed API-heavy systems. 5) Bearer Tokens and JWT # These two terms are related, but they are not the same.\nA bearer token means:\nwhoever holds the token can use it\nA JWT means:\na token formatted as a compact JSON-based structure with claims inside it\nSo:\nBearer describes how the token is used JWT describes what the token looks like A bearer token can be a JWT. But not every bearer token is a JWT.\nWhy they exist # Modern systems, especially APIs, often do not want to use server sessions for everything.\nInstead, they want the client to send an access credential directly with each request.\nThat is where bearer tokens help.\nJWT becomes popular because it can carry useful claims inside the token, such as:\nwho the user is who issued the token who the audience is when it expires Flow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant Client participant AuthServer participant API Client-\u003e\u003eAuthServer: Authenticate / obtain token AuthServer--\u003e\u003eClient: Token Client-\u003e\u003eAPI: Request + Bearer token API-\u003e\u003eAPI: Validate token alt Valid API--\u003e\u003eClient: Protected response else Invalid API--\u003e\u003eClient: Unauthorized end Example shape # Bearer header:\nAuthorization: Bearer eyJhbGciOi... JWT shape:\nheader.payload.signature Example payload shape:\n{ \u0026#34;sub\u0026#34;: \u0026#34;user123\u0026#34;, \u0026#34;iss\u0026#34;: \u0026#34;auth.example.com\u0026#34;, \u0026#34;aud\u0026#34;: \u0026#34;api.example.com\u0026#34;, \u0026#34;exp\u0026#34;: 1717000000 } Main objects / terminologies # Bearer token A token usable by whoever possesses it.\nJWT A structured token format carrying claims.\nClaims Data inside the token, like user ID or expiration time.\nSignature Protects integrity, meaning tampering can be detected.\nIssuer (iss) Who created the token.\nAudience (aud) Who the token is meant for.\nExpiration (exp) When the token is no longer valid.\nCaveats # If a bearer token is stolen, it can often be used. JWT payloads are encoded, not automatically encrypted. Never trust a JWT just because you can decode it. The token must be validated properly. Bearer and JWT should not be treated as synonyms. 6) Access Tokens and Refresh Tokens # This topic answers a very practical question:\nWhy not just issue one token and use it forever?\nBecause that would be bad for security.\nSo modern systems often split the job into two tokens:\nAccess token: used to call the API Refresh token: used to get a new access token This creates a balance:\nshort-lived working token for normal use longer-lived renewal token for user convenience Why it exists # If access tokens last too long, stolen tokens remain useful for too long.\nIf access tokens expire too quickly without renewal, users must log in too often.\nRefresh tokens solve that tradeoff.\nSo the design goal is:\nbetter security without terrible user experience\nFlow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant Client participant AuthServer participant API Client-\u003e\u003eAuthServer: Login / token request AuthServer--\u003e\u003eClient: Access token + Refresh token Client-\u003e\u003eAPI: Use access token API--\u003e\u003eClient: Response Client-\u003e\u003eAPI: Use expired access token API--\u003e\u003eClient: Unauthorized Client-\u003e\u003eAuthServer: Send refresh token AuthServer--\u003e\u003eClient: New access token Client-\u003e\u003eAPI: Retry with new access token API--\u003e\u003eClient: Response Example shape # Access token usage:\nAuthorization: Bearer \u0026lt;access_token\u0026gt; Refresh request shape:\nPOST /token Content-Type: application/x-www-form-urlencoded grant_type=refresh_token\u0026amp;refresh_token=\u0026lt;refresh_token\u0026gt; Main objects / terminologies # Access token The token used to access protected APIs.\nRefresh token The token used to obtain a new access token.\nExpiration Access tokens usually expire sooner.\nRotation Replacing an old refresh token with a new one after use.\nCaveats # Refresh tokens are usually more sensitive than people think. Refresh tokens should not be sent to normal APIs. Access tokens are often short-lived by design. Refresh tokens may also expire or be revoked. Rotation is often used to reduce abuse risk. 7) OAuth 2.0 # OAuth 2.0 is not mainly “login.” It is a framework for delegated authorization.\nThat means:\none app can get limited permission to access something on behalf of a user\nwithout asking the user to give that app their password directly.\nThis is the core problem OAuth solves.\nExample:\nYou want a calendar app to access your Google Calendar. The app redirects you to Google. You approve access there. The app gets tokens instead of your Google password. That is OAuth thinking.\nWhy it exists # Without OAuth, third-party apps might ask users for their passwords directly.\nThat is dangerous and hard to control.\nOAuth improves this by introducing:\npermission screens scopes limited tokens central authorization So the user can say:\n“I allow this app to do this specific thing”\ninstead of:\n“Here is my password.”\nFlow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant User participant ClientApp participant AuthServer participant ResourceServer ClientApp-\u003e\u003eUser: Ask to connect account User-\u003e\u003eAuthServer: Login and approve access AuthServer--\u003e\u003eClientApp: Authorization result / code ClientApp-\u003e\u003eAuthServer: Exchange for token AuthServer--\u003e\u003eClientApp: Access token ClientApp-\u003e\u003eResourceServer: Call API with token ResourceServer--\u003e\u003eClientApp: Protected resource Example shape # Authorization request shape:\nGET /authorize? response_type=code \u0026amp;client_id=client123 \u0026amp;redirect_uri=https://app.example.com/callback \u0026amp;scope=read_profile \u0026amp;state=xyz Token exchange shape:\nPOST /token grant_type=authorization_code \u0026amp;code=abc123 \u0026amp;redirect_uri=https://app.example.com/callback Main objects / terminologies # Resource owner Usually the user who owns the data.\nClient The application asking for access.\nAuthorization server The system that authenticates the user and grants tokens.\nResource server The API holding the protected data.\nScope What level of access is being requested.\nAuthorization code A short-lived code later exchanged for tokens.\nPKCE A protection added to modern authorization code flows.\nCaveats # OAuth is about authorization, not identity by itself. OAuth does not automatically tell the client who the user is in a standardized way. Older flows are less preferred now. Redirect security matters a lot. Best understood as permission delegation. 8) OpenID Connect # OpenID Connect, or OIDC, is the identity layer added on top of OAuth 2.0.\nOAuth answers:\n“Can this app access this resource?”\nOIDC answers:\n“Who is the user who just signed in?”\nThat is the main job of OIDC.\nIt adds a special token called the ID Token, which tells the client about the user’s authentication.\nWhy it exists # OAuth alone is not enough for modern login.\nAn application often needs to know:\nwho signed in which provider authenticated them whether the sign-in result is trustworthy OIDC standardizes that.\nSo instead of every provider inventing its own login identity format, OIDC gives a shared model.\nFlow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant User participant App participant IdentityProvider participant UserInfo App-\u003e\u003eIdentityProvider: Login request with openid scope User-\u003e\u003eIdentityProvider: Authenticate IdentityProvider--\u003e\u003eApp: ID Token + Access Token App-\u003e\u003eApp: Validate ID Token App-\u003e\u003eUserInfo: Optional request for more profile data UserInfo--\u003e\u003eApp: User claims Example shape # Authorization request:\nGET /authorize? response_type=code \u0026amp;client_id=client123 \u0026amp;redirect_uri=https://app.example.com/callback \u0026amp;scope=openid profile email \u0026amp;state=xyz \u0026amp;nonce=abc ID token payload shape:\n{ \u0026#34;iss\u0026#34;: \u0026#34;https://idp.example.com\u0026#34;, \u0026#34;sub\u0026#34;: \u0026#34;user123\u0026#34;, \u0026#34;aud\u0026#34;: \u0026#34;client123\u0026#34;, \u0026#34;exp\u0026#34;: 1717000000, \u0026#34;iat\u0026#34;: 1716996400 } Main objects / terminologies # OIDC Identity layer on top of OAuth 2.0.\nID Token A token telling the client about the authentication result and user identity.\nopenid scope The switch that makes the request OIDC.\nUserInfo endpoint An endpoint used to fetch more user claims.\nNonce A value used to tie the request and token together more safely.\nClaims Identity data such as subject, email, or name.\nCaveats # OIDC does not replace OAuth; it builds on top of it. The ID token is not the same as the access token. APIs usually want access tokens, not ID tokens. ID tokens must be validated properly. OIDC is the right model when the app needs standardized login identity. 9) Single Sign-On (SSO) # Single Sign-On means:\nsign in once, access many applications\nSSO is not one specific protocol. It is the overall login experience created when multiple apps trust the same identity provider.\nSo SSO is the user-facing result.\nWhy it exists # Without SSO, users sign in separately to:\nemail HR portal dashboard internal tools support systems That is annoying and hard to manage.\nSSO improves this by centralizing login.\nOne identity system handles authentication, and many apps trust it.\nFlow # --- config: theme: 'base' themeVariables: primaryColor: '#BB2528' primaryTextColor: '#0f0e0e' --- sequenceDiagram participant User participant AppA participant AppB participant IdentityProvider User-\u003e\u003eAppA: Open app A AppA-\u003e\u003eIdentityProvider: Redirect for login User-\u003e\u003eIdentityProvider: Sign in IdentityProvider--\u003e\u003eAppA: Successful authentication User-\u003e\u003eAppB: Open app B AppB-\u003e\u003eIdentityProvider: Redirect for login IdentityProvider-\u003e\u003eIdentityProvider: Existing login session found IdentityProvider--\u003e\u003eAppB: Successful authentication Example shape # Conceptually:\nApp A trusts the identity provider App B trusts the same identity provider user signs in once at the identity provider both apps accept that result SSO is often implemented using:\nOIDC SAML Main objects / terminologies # SSO Single Sign-On, one login used across multiple apps.\nIdentity Provider (IdP) The central system that authenticates the user.\nRelying Party / Service Provider The application that trusts the identity provider.\nFederation Trust between separate systems for identity.\nSingle Logout (SLO) A separate idea: signing out across apps.\nCaveats # SSO is not the same as one app session. SSO is not the same thing as OAuth by itself. SSO does not automatically mean global logout. If the central identity provider is compromised, many apps may be affected. OIDC and SAML are common ways to implement SSO. Final comparison: how to think about all of them # Here is the simplest mental map.\nDirect credential methods # Basic Auth: send username/password directly Digest Auth: prove knowledge of password with a challenge Client/app identification # API Key: identify the calling application or integration App-managed login # Session-based auth: server remembers the logged-in user Token-based API access # Bearer token: whoever has it can use it JWT: a common format for token contents Access token: used to call APIs Refresh token: used to get a new access token Delegated authorization and identity # OAuth 2.0: app gets permission to access resources OIDC: app learns who the user is SSO: one login works across many apps The big picture # If you remember only one thing, remember this:\nBasic / Digest are old-style ways to prove credentials on requests. API keys identify apps more than users. Sessions let the server remember a logged-in user. Bearer/JWT are modern token ideas for APIs. Access + Refresh tokens balance security and usability. OAuth 2.0 is about permission delegation. OIDC is about identity. SSO is about one login across many apps. References # How to Design APIs (REST, GraphQL, Auth, Security) - Hayk Simonyan ","date":"29 April 2026","externalUrl":null,"permalink":"/blogs/authentication/","section":"Blogs","summary":"","title":"Authentication-Explained : From Basic Auth to SSO","type":"blogs"},{"content":"","date":"29 April 2026","externalUrl":null,"permalink":"/blogs/","section":"Blogs","summary":"","title":"Blogs","type":"blogs"},{"content":"When engineers first approach DynamoDB, the biggest mistake is treating it like a faster version of SQL.\nIt is not.\nDynamoDB is a fully managed NoSQL database built for fast, predictable performance at massive scale. But to use it well, you have to change the way you think about data. In a relational database, you typically design normalized tables first and let SQL answer many kinds of questions later. In DynamoDB, you usually do the opposite: you start with the queries your application needs, and then design your data model around those access patterns.\nThat shift is what makes DynamoDB both powerful and tricky. Used well, it can outperform traditional relational systems for the right workloads. Used poorly, it can become expensive, awkward, and hard to evolve.\nThe DynamoDB mindset # At the heart of DynamoDB is the idea that your primary key is not just an identifier — it is the foundation of how your data is stored, distributed, and retrieved.\nA DynamoDB table uses either a simple primary key made of a partition key, or a composite primary key made of a partition key and sort key. The partition key determines how data is distributed internally, and the sort key lets you organize related items within the same partition. This makes DynamoDB especially strong when your application naturally asks questions like: “get all orders for this customer,” “get the latest messages in this chat,” or “get all events for this device in time order.”\nThis is why DynamoDB design is really query-first design.\nInstead of asking, “What tables should I create?”, the better question is:\nWhat are the exact reads and writes my application performs most often?\nDesigning for queries, not relationships # In SQL, if you have customers and orders, you would typically create separate tables and join them when needed.\nIn DynamoDB, a better design is often to keep related records together under the same partition key, so they can be fetched efficiently with a single query. DynamoDB calls all items sharing the same partition key an item collection. This pattern works well for one-to-many relationships.\nImagine an e-commerce system where one of the main features is:\nshow all orders for a customer show recent orders for a customer show orders for a customer within a date range A DynamoDB-friendly model could look like this:\nPK = CUSTOMER#123 SK = ORDER#2026-04-26T10:15:00Z#ORD789 With this structure, all of a customer’s orders live under the same partition key, and the sort key keeps them in time order. That makes queries efficient and natural:\nall orders for a customer latest N orders orders between two dates This is a good example of a data model that fits DynamoDB’s architecture because the access pattern is known up front and the sort key does useful work.\nQuery is the happy path. Scan is the warning sign. # One of the most important DynamoDB concepts is the difference between Query and Scan.\nA Query is efficient because it targets a specific partition key and can optionally narrow results using the sort key. A Scan, on the other hand, reads every item in a table or index. That means scans are often slow, expensive, and a sign that the data model does not match the application’s real access patterns. AWS’s own guidance is clear here: scans read all items, and even filter expressions are applied after data is read.\nThat last point matters a lot.\nMany developers assume this kind of query is efficient:\n“Read everything for a customer, then filter to only OPEN items.”\nBut if you read a large set and discard most of it afterward, you still pay for what was read. In DynamoDB, filtering does not rescue a weak primary-key design.\nSecondary indexes: where GSI and LSI come in # Sooner or later, one table key is not enough.\nYou may model orders by customer, but then the product team asks for:\nshow all open tickets assigned to Alice show all orders by status show all products by category show all messages sorted a different way That is where secondary indexes enter the picture.\nGlobal Secondary Index (GSI) # A GSI gives you a completely different access path to your data. It can use a different partition key and a different sort key from the base table. DynamoDB keeps the index synchronized automatically, but updates happen asynchronously, so GSI reads are eventually consistent.\nThis makes GSIs ideal when you need to query the same data from a totally different angle.\nFor example, suppose the base table stores tickets like this:\nPK = TENANT#42 SK = TICKET#9001 But your product also needs:\nall OPEN tickets assigned to Alice You could add a GSI like this:\nGSI1PK = ASSIGNEE#alice GSI1SK = STATUS#OPEN#2026-04-26T08:00:00Z Now DynamoDB can answer a query that the base table was never designed for.\nThat is the real value of a GSI: it gives you a new, queryable view of the same underlying items.\nLocal Secondary Index (LSI) # An LSI, or local secondary index, is more constrained but useful in specific cases. It keeps the same partition key as the base table and only changes the sort key. LSIs let you query the same parent group in different sort orders or dimensions. They support strongly consistent reads, but they must be created when the table is created. Also, each item collection for a given partition key cannot exceed 10 GB when LSIs are involved.\nA good LSI use case is when the “parent” always stays the same, but you want multiple ways to sort that parent’s items.\nFor example, for a bank account you might want:\ntransactions ordered by event time transactions ordered by settlement date transactions ordered by amount If the partition key is always ACCOUNT#555, then an LSI is a natural fit because the account remains the same and only the sort dimension changes.\nThe simplest way to choose # A practical rule is:\nuse an LSI when you want a different sort order within the same partition key use a GSI when you need a totally different lookup path across the table That one distinction prevents a lot of confusion.\nWhere DynamoDB fits naturally # DynamoDB performs best when your workload has a few characteristics:\nvery high request volume predictable and well-known access patterns key-based reads and writes low-latency requirements limited need for joins clear scaling expectations This is why DynamoDB is commonly a great fit for things like session stores, shopping carts, device state, user preferences, game state, activity feeds, and time-ordered event streams. These systems usually care more about fast, predictable reads and writes than about complex relational queries. AWS describes DynamoDB as a fully managed NoSQL service with fast and predictable performance and seamless scalability, which lines up directly with these kinds of workloads.\nA concrete case where DynamoDB can outperform SQL # A good example is a session store for a high-traffic consumer application.\nImagine millions of users opening an app and repeatedly doing simple operations like:\nget session by ID update cart quantity refresh expiration fetch current user state This is close to DynamoDB’s ideal workload. The access pattern is simple, the keys are known, the reads and writes are small, and the application benefits from predictable low latency at scale. A relational database can absolutely support this, but doing so often requires more operational work around partitioning, caching, connection scaling, and query tuning. DynamoDB is built to absorb that pattern more directly.\nThat does not mean DynamoDB is “better than SQL” in general.\nIt means DynamoDB is often better than SQL for workloads that are mostly high-scale key lookups and targeted writes, especially when joins are not central to the product.\nArchitecture implications: why the model matters so much # Because DynamoDB distributes data using the partition key, the partition key is also your scaling strategy.\nIf too much traffic hits a small number of partition key values, you can create hot partitions. This is why a good key design needs enough cardinality and traffic spread. It is also why DynamoDB modeling is deeply tied to the architecture of the system, not just the shape of the data.\nThis is different from SQL thinking. In relational design, schema structure is often separate from physical scaling decisions. In DynamoDB, the two are tightly connected.\nThat is one reason DynamoDB shows up so often in system design interviews: it forces you to think about storage layout, query patterns, and scale together.\nAdvanced features that make DynamoDB more practical # DynamoDB is not just a simple key-value store. It also includes features that make it more usable in real systems.\nIt supports transactions for multi-item ACID operations, though with specific limits: up to 100 items and 4 MB total per transaction. It also supports expressions for conditional updates, filtering, and update logic. These features help when you need correctness beyond a single item, but they do not turn DynamoDB into a full relational engine.\nThat distinction matters. Transactions help with correctness. They do not solve data modeling mistakes.\nThe biggest limitations of DynamoDB # DynamoDB is powerful, but it comes with real tradeoffs.\nThe first is that it is much less friendly to ad hoc querying than SQL. If a new business requirement appears and your table was not designed for that access pattern, you may need a new GSI, additional denormalization, or even a redesign.\nThe second is that DynamoDB does not support relational joins in the way SQL databases do. If your application naturally depends on joining many entities in flexible ways, a relational database will often be a better fit.\nThe third is that bad key design can create operational pain. If a small number of partition keys absorb too much traffic, performance becomes uneven and the benefits of DynamoDB’s scale are harder to realize.\nThere are also hard constraints to respect. DynamoDB has a maximum item size of 400 KB, which means it is not a good place for large blobs or oversized documents. When items exceed that size, AWS recommends patterns like splitting data across items or storing large payloads in S3 and saving references in DynamoDB.\nAnd while indexes are powerful, they add tradeoffs too. GSIs introduce additional write and storage overhead and are eventually consistent. LSIs must be defined up front and come with the 10 GB item-collection limit per partition key.\nSo when should you use DynamoDB? # Use DynamoDB when:\nyour access patterns are known in advance you need high throughput and low latency you are comfortable modeling data around queries your workload is dominated by key-based access you want a managed system that scales cleanly Use SQL when:\nyour queries evolve frequently joins are central to the application reporting and relational integrity are core requirements the shape of the data matters less than the flexibility of the query engine The real lesson is not “DynamoDB vs SQL.”\nThe real lesson is that they solve different kinds of problems.\nFinal thought # The best way to understand DynamoDB is to stop asking, “How do I map my SQL schema into DynamoDB?”\nInstead, ask:\nWhat are the exact reads and writes my system performs, at what scale, and how can I model my keys so those operations are efficient by default?\nThat is the DynamoDB mindset.\nAnd once that clicks, concepts like partition keys, sort keys, GSIs, LSIs, and query-driven modeling stop feeling like limitations. They start feeling like tools for building systems that are simple, fast, and scalable by design.\nReference # Hello Interview AWS - Documentation ","date":"26 April 2026","externalUrl":null,"permalink":"/blogs/dynamo-db/","section":"Blogs","summary":"","title":"DynamoDB, Explained: How to Think in Access Patterns Instead of Tables","type":"blogs"},{"content":"If you are building modern, high-performance distributed systems, caching isn\u0026rsquo;t just an optimization—it is a foundational requirement. At the center of this ecosystem sits Redis. Often misunderstood as \u0026ldquo;just a cache,\u0026rdquo; Redis is actually a highly versatile, in-memory data structure server.\nWhether you are designing a real-time gaming leaderboard, protecting your APIs with rate limiters, or scaling a high-throughput microservice architecture, mastering Redis is a superpower. Let\u0026rsquo;s dive deep into how it works, the patterns that define it, and how to scale it in production.\n1. What is Redis and Why Redis? # Redis (Remote Dictionary Server) is an open-source, in-memory, key-value data store. Unlike traditional relational databases (like PostgreSQL or MySQL) that write data to spinning disks or SSDs, Redis holds all of its data in RAM.\nWhy choose Redis?\nBlistering Speed: Because accessing RAM is orders of magnitude faster than accessing a disk, Redis routinely delivers sub-millisecond response times, handling millions of operations per second. Rich Data Structures: It doesn’t just store dumb strings. It natively understands complex data structures like Lists, Sets, Hashes, and Sorted Sets, allowing you to offload computational complexity from your application servers to the database. Versatility: It operates effectively as a database, a cache, a message broker, and a streaming engine. 2. Redis Architecture: The Secret to Speed # Understanding Redis requires understanding its somewhat counterintuitive architectural choices.\nThe Single-Threaded Event Loop # The most surprising fact about Redis is that it uses a single thread to process commands. In an era of multi-core processors, why limit a database to one thread? Because memory access is so fast, the CPU is almost never the bottleneck—network I/O and memory bandwidth are. By using a single thread, Redis completely eliminates the overhead of context switching, thread locks, and race conditions. It handles massive concurrency using an I/O multiplexing model (like epoll), queuing up thousands of requests and executing them sequentially with ruthless efficiency.\nPersistence Mechanisms # Even though Redis is \u0026ldquo;in-memory,\u0026rdquo; it doesn\u0026rsquo;t mean your data vanishes when the server reboots. Redis offers robust persistence:\nRDB (Redis Database): Takes point-in-time snapshots of your data and saves them to disk at specified intervals. AOF (Append Only File): Logs every single write operation received by the server. If the server crashes, Redis replays this log to reconstruct the exact state. 3. Data Structures in Action # The brilliance of Redis lies in its data structures. Instead of querying a relational database and sorting data in your application code, you push data into a structure that is already optimized for what you need to do.\nA. Strings # The most basic Redis type. A string can contain text, serialized objects (like JSON), or even binary data (up to 512MB).\nApplication Scenario: Session caching, HTML page caching, and rate-limiting/counters. Example (Counters): Because Redis is single-threaded, commands like INCR (increment) are atomic. You can use it to count page views or API requests safely without race conditions. SET api_requests:user1001 0 INCR api_requests:user1001 // Returns 1 B. Hashes # Hashes are maps between string fields and string values. Think of them as flat JSON objects or a row in a SQL database.\nApplication Scenario: Storing object data, like user profiles, where you frequently need to access or update individual fields rather than the whole object. Example (User Profile): HSET user:205 name \u0026#34;Alice\u0026#34; role \u0026#34;Admin\u0026#34; status \u0026#34;Active\u0026#34; HGET user:205 role // Returns \u0026#34;Admin\u0026#34; C. Lists # Lists are linked lists of strings. Because they are linked lists, inserting at the head or tail is incredibly fast — $O(1)$ complexity — but accessing elements by index in the middle is slower — $O(N)$.\nApplication Scenario: Message queues, activity streams, or \u0026ldquo;latest N items\u0026rdquo; (like a Twitter timeline). Example (Activity Stream): Pushing a new activity to the front of a list and keeping only the latest 100. LPUSH timeline:alice \u0026#34;tweet_id_99\u0026#34; LTRIM timeline:alice 0 99 // Trims the list to only keep the newest 100 D. Sets # Sets are unordered collections of unique strings. You can easily add, remove, and test for the existence of members in $O(1)$ time.\nApplication Scenario: Tracking unique items (e.g., unique IP addresses visiting a site), tagging systems, or finding relationships (intersections/unions). Example (Mutual Friends): SADD friends:alice \u0026#34;bob\u0026#34; \u0026#34;charlie\u0026#34; \u0026#34;david\u0026#34; SADD friends:eve \u0026#34;charlie\u0026#34; \u0026#34;david\u0026#34; \u0026#34;frank\u0026#34; SINTER friends:alice friends:eve // Returns \u0026#34;charlie\u0026#34;, \u0026#34;david\u0026#34; (mutual friends) E. Sorted Sets (ZSets) # Similar to Sets, but every element is associated with a floating-point number called a \u0026ldquo;score.\u0026rdquo; Elements are always kept ordered by this score.\nApplication Scenario: Leaderboards, priority queues, time-series data (using timestamps as the score). Example (Gaming Leaderboard): ZADD global_leaderboard 1500 \u0026#34;PlayerA\u0026#34; ZADD global_leaderboard 2100 \u0026#34;PlayerB\u0026#34; ZREVRANGE global_leaderboard 0 9 WITHSCORES // Gets top 10 players sorted highest to lowest Examples # 1. The Rate Limiter (API Protection) # When you have a popular API, you must protect your backend databases from being overwhelmed by too many requests from a single user or IP address.\nThe Data Structure: Strings (for Fixed Window) or Sorted Sets (for Sliding Window). Let\u0026rsquo;s look at the most robust method for exact accuracy: the Sliding Window Log using Sorted Sets (ZSET).\nHow it works: We use the timestamp of the request as both the score and the value in a Sorted Set. When a request comes in, we remove all timestamps older than our window, count what is left, and decide if the request is allowed.\nRedis Commands (Simulating a limit of 3 requests per 60 seconds for User 123):\n// 1. A request comes in at timestamp 1700000000. Add it to the set. ZADD rate_limit:user123 1700000000 \u0026#34;1700000000\u0026#34; // 2. Remove any requests that happened more than 60 seconds ago (score \u0026lt; 1699999940) ZREMRANGEBYSCORE rate_limit:user123 -inf 1699999940 // 3. Count how many requests are in the current window ZCARD rate_limit:user123 // 4. Set an expiry on the whole key so we don\u0026#39;t leak memory for inactive users EXPIRE rate_limit:user123 60 Architectural Note: In a real application, you would execute these four commands together inside a Redis Lua script so they run atomically, preventing race conditions between concurrent requests.\n2. The Distributed Lock (Resource Synchronization) # Imagine you have three instances of a worker service, and they all wake up at midnight to process the exact same daily billing report. If they all run it, you charge customers three times. You need a lock that spans across your distributed system.\nThe Data Structure: Strings.\nHow it works: You attempt to set a key. If the key already exists, someone else has the lock. If it doesn\u0026rsquo;t, you get the lock. Crucially, you must set an expiration time (Time-To-Live or TTL) so that if your worker crashes while holding the lock, the lock eventually releases itself.\nRedis Commands:\n// Attempt to acquire the lock. // NX = Only set if it does NOT exist. // PX 30000 = Expire in 30,000 milliseconds (30 seconds). SET lock:billing_report \u0026#34;worker_node_1_random_uuid\u0026#34; NX PX 30000 If Redis returns OK, this worker has the lock. If it returns (nil), another worker has it, and this worker should back off.\nArchitectural Note: Why the random UUID? When a worker finishes, it needs to delete the lock. However, if the worker took too long and the lock auto-expired, another worker might have grabbed it. The worker must check if the value matches its own UUID before deleting it, ensuring it doesn\u0026rsquo;t accidentally delete someone else\u0026rsquo;s lock.\n3. The Real-Time Leaderboard (Gaming or Analytics) # Relational databases struggle with massive, real-time leaderboards. Running ORDER BY score DESC on a million rows every time a user checks their rank is a recipe for database failure.\nThe Data Structure: Sorted Sets (ZSET).\nHow it works: Sorted sets maintain data in memory already perfectly ordered by the score. Fetching the top 10, or finding a specific user\u0026rsquo;s rank, is an $O(\\log(N))$ operation.\nRedis Commands:\n// Add players and their scores ZADD global_leaderboard 1500 \u0026#34;PlayerA\u0026#34; ZADD global_leaderboard 2100 \u0026#34;PlayerB\u0026#34; ZADD global_leaderboard 800 \u0026#34;PlayerC\u0026#34; // Player A wins a match and gains 50 points ZINCRBY global_leaderboard 50 \u0026#34;PlayerA\u0026#34; // Get the top 2 players (highest score first) ZREVRANGE global_leaderboard 0 1 WITHSCORES // Find Player C\u0026#39;s exact rank (0-indexed, starting from highest score) ZREVRANK global_leaderboard \u0026#34;PlayerC\u0026#34; 4. Geo-Spatial Queries (Ride-sharing or Proximity Search) # If you are building an app that needs to find \u0026ldquo;drivers within a 5km radius of the user,\u0026rdquo; calculating the Haversine formula across a SQL database of coordinates is intensely slow.\nThe Data Structure: Geo Maps (which are actually Sorted Sets under the hood, utilizing a geohash algorithm to map 2D coordinates into a 1D score).\nHow it works: You add members with their longitude and latitude. Redis handles the complex math to index them and allows you to search by radius or bounding box.\nRedis Commands:\n// Add drivers to the map (Longitude, Latitude, Name) GEOADD drivers -122.0289 37.3323 \u0026#34;Driver_Alice\u0026#34; GEOADD drivers -121.8863 37.3382 \u0026#34;Driver_Bob\u0026#34; // Located in San Jose // Find all drivers within 10 kilometers of a user\u0026#39;s coordinate GEOSEARCH drivers FROMLONLAT -121.8900 37.3300 BYRADIUS 10 km WITHDIST 4. Redis Atomicity, Lua Scripts, and Python # Because Redis is single-threaded, individual commands are atomic. However, what if you need to run multiple commands together safely, ensuring no other client interrupts them?\nYou could use Redis Transactions (MULTI/EXEC), but Lua Scripts are the modern industry standard. When you send a Lua script to Redis, the entire script is executed atomically. This prevents race conditions and drastically reduces network latency by combining multiple operations into a single round-trip.\nPython redis-py Lua Script Example: A Flawless Rate Limiter # Here is how you execute an atomic rate limiter in Python using a Lua script. We use register_script so the script is compiled and cached on the Redis server, saving bandwidth.\nimport redis # Connect to Redis r = redis.Redis(host=\u0026#39;localhost\u0026#39;, port=6379, decode_responses=True) # 1. Define the Lua Script # Increments a key, sets a TTL if it\u0026#39;s new, and returns 1 if allowed, 0 if blocked. lua_script = \u0026#34;\u0026#34;\u0026#34; local current = redis.call(\u0026#34;INCR\u0026#34;, KEYS[1]) if current == 1 then redis.call(\u0026#34;EXPIRE\u0026#34;, KEYS[1], ARGV[1]) end if current \u0026gt; tonumber(ARGV[2]) then return 0 end return 1 \u0026#34;\u0026#34;\u0026#34; # 2. Register the script with the Redis server rate_limit_script = r.register_script(lua_script) def is_allowed(user_id, window_seconds=60, max_requests=10): key = f\u0026#34;rate:{user_id}\u0026#34; # 3. Execute the script atomically return bool(rate_limit_script(keys=[key], args=[window_seconds, max_requests])) # Usage if is_allowed(\u0026#34;user_999\u0026#34;): print(\u0026#34;Request processed!\u0026#34;) else: print(\u0026#34;HTTP 429: Too Many Requests\u0026#34;) 5. Caching Patterns: When and Why # Moving from how to write Redis commands to where to place Redis in your overall architecture is what separates good developers from great system designers.\nWhen we talk about caching patterns, we are defining the relationship and the flow of data between three core components: your Application, the Cache (Redis), and the Primary Database (like PostgreSQL or MongoDB).\n1. Cache-Aside (Lazy Loading) # How it works (Read): App asks Cache. If miss, App asks Database. App saves to Cache. App returns data. How it works (Write): App writes directly to Database. App then deletes (invalidates) the specific key in the Cache. Reasoning: It is highly resilient. If Redis goes down, your application still works (it just hits the database directly and gets slower). It\u0026rsquo;s perfect for read-heavy workloads. By only caching what is actually requested (lazy loading), you don\u0026rsquo;t waste memory on unused data. Example: Loading a user\u0026rsquo;s profile. You don\u0026rsquo;t need to load every user into Redis on startup; you only cache the profiles of users who are currently active. 2. Read-Through # In this pattern, the application treats the cache as the main data store. The application code never talks to the database directly for reads.\nHow it works (Read): App asks Cache. If miss, the Cache itself (or a smart caching library bridging them) fetches from the Database, updates itself, and returns the data to the App. Reasoning: It dramatically simplifies application code. Your app just says get(user_id) and doesn\u0026rsquo;t care where it comes from. It ensures the cache and database schemas are aligned. Example: A Content Management System (CMS) like WordPress. The application asks for the article content. A data access layer handles checking Redis, fetching from MySQL if needed, and returning the result. 3. Write-Through # This pattern prioritizes absolute data consistency over write speed.\nHow it works (Write): App writes to the Cache. The Cache immediately and synchronously writes to the Database. Only when both are successful does the operation return to the App. Reasoning: You use this when you cannot tolerate stale data under any circumstances. Every read against the cache will always reflect the absolute latest state of the database. The tradeoff is that writes are slower because they incur the latency of two network hops (App -\u0026gt; Cache -\u0026gt; DB). Example: Banking transactions or real-time inventory systems where selling an item that isn\u0026rsquo;t actually in stock is catastrophic. 4. Write-Behind (Write-Back) # This is the most aggressive pattern for write performance, but it comes with the highest risk.\nHow it works (Write): App writes data to the Cache. The Cache immediately returns \u0026ldquo;Success\u0026rdquo; to the App. Asynchronously, in the background, the Cache batches those updates and flushes them to the Database later. Reasoning: Extreme performance. Your application\u0026rsquo;s write speed is only limited by Redis\u0026rsquo;s RAM speed. It completely shields your slow database from massive spikes in write traffic. However, if the Redis server crashes before the background flush occurs, that data is permanently lost. Example: Counting \u0026ldquo;Likes\u0026rdquo; on a viral social media post, capturing high-velocity IoT sensor data, or tracking video view progress. If you lose a few milliseconds of \u0026ldquo;Likes\u0026rdquo; during a crash, it\u0026rsquo;s not the end of the world. 5. Write-Around # This pattern is used to optimize the cache\u0026rsquo;s memory space and prevent \u0026ldquo;cache pollution.\u0026rdquo;\nHow it works (Write): App writes directly to the Database, completely bypassing the Cache. Reasoning: You use this when data is written once but rarely read back immediately. If you write massive log files or upload large images using Write-Through, you fill up your expensive Redis RAM with data nobody is looking at, kicking out important data. Write-Around ensures only data that is explicitly requested gets cached. Example: Uploading an archival document, chat history backups, or generating monthly PDF reports. Example scenarios in a CRM product # 1. Categorizing CRM Caching Requirements # To design the right strategy, we must first segment the data based on its Read/Write ratio and Consistency requirements:\nStatic Metadata \u0026amp; Config: (e.g., Custom field definitions, dropdown values, UI layouts). Requirement: High Read, Very Low Write. Must be available globally. Pattern: Read-Through. User Sessions \u0026amp; Permissions: (e.g., Auth tokens, RBAC roles). Requirement: Extremely High Read (checked on every API call). Needs fast expiration. Pattern: Distributed Session Store (Redis Hash). Customer Entities: (e.g., Contact details, Lead info). Requirement: Read-heavy but requires \u0026ldquo;Read-Your-Own-Writes\u0026rdquo; consistency. Pattern: Cache-Aside with explicit invalidation. Activity Streams \u0026amp; Interaction Logs: (e.g., \u0026ldquo;User A called Lead B\u0026rdquo;). Requirement: Write-heavy. Absolute real-time consistency is often less critical than ingestion speed. Pattern: Write-Behind (Write-Back). Analytics \u0026amp; Pipeline Reports: (e.g., \u0026ldquo;Total sales forecast for Q3\u0026rdquo;). Requirement: Computationally expensive queries. Data changes frequently, but users tolerate \u0026ldquo;stale\u0026rdquo; data for a few minutes. Pattern: Timed Invalidation (TTL-based). 2. Solving Requirements with Specific Patterns # Scenario A: High-Velocity Activity Logs (The Throughput Problem) # In a large CRM, thousands of automated syncs (from email, Slack, or VoIP) hit the system simultaneously. Writing every single log to a relational database (SQL) immediately will cause lock contention.\nSolution: Use Write-Behind. The application writes the activity to a Redis List or Stream. A background worker batches these and writes them to the database every 5 seconds. This flattens the spike and protects the DB. Scenario B: Field Permissions \u0026amp; Metadata (The Consistency Problem) # If an admin changes a user\u0026rsquo;s permission from \u0026ldquo;Editor\u0026rdquo; to \u0026ldquo;Viewer,\u0026rdquo; that change must be reflected immediately across all app instances to prevent unauthorized edits.\nSolution: Use Cache-Aside with Invalidation. When the admin saves the change (Write-Around to DB), the application explicitly issues a DEL command to the Redis key. The next request will result in a cache miss, forcing a fresh pull of the new permissions. Scenario C: Sales Dashboards (The Latency Problem) # Calculating a \u0026ldquo;Sales Pipeline\u0026rdquo; involves joining multiple tables (Leads, Deals, Quotes, Users). Running this on every page load is too slow.\nSolution: Pre-computation. Instead of calculating on demand, a scheduled job calculates the dashboard data every 10 minutes and stores the final JSON in Redis. 3. Advanced Considerations: The \u0026ldquo;Thundering Herd\u0026rdquo; in CRM # In a CRM, when a major sales report expires in the cache, multiple managers might refresh their browser at the exact same second. If the cache is empty, all those requests hit the database simultaneously to re-calculate the same report.\nSolution: Locking/Single-Flighting. When the first request sees a cache miss, it acquires a Redis Distributed Lock. Subsequent requests see the lock is held and simply wait for the first process to finish updating the cache, or they serve the \u0026ldquo;stale\u0026rdquo; version temporarily. This ensures the expensive database query only runs once.\n6. Cache Eviction Policies # If you buy a 2GB Redis server, what happens when you try to write 2.1GB of data? By default, Redis throws an Out of Memory (OOM) error and halts all writes.\nIf you are using Redis purely as a cache, you want it to delete old data to make room for new data. You configure this using Eviction Policies:\nallkeys-lru (Least Recently Used): The most popular setting. Evicts the keys that haven\u0026rsquo;t been accessed in the longest amount of time. allkeys-lfu (Least Frequently Used): Evicts keys that are rarely accessed overall, protecting viral content that might have a brief pause in traffic. volatile-ttl: Only evicts keys that have an expiration set (TTL), prioritizing the deletion of keys that are closest to expiring anyway. 7. Scenario-Based Cost and Memory Calculations # Because RAM is significantly more expensive than disk storage (averaging $15-$20 per GB on managed cloud providers), memory planning is critical. You must account for Redis Overhead—the memory Redis uses to store the key name, expiration timers, and internal pointers.\nScenario A: High-Volume Session Store (Hashes)\nData: 1,000,000 active user sessions. Payload: 500 bytes of data per session. Math: Raw Data (500 MB) + Hash Overhead (~90 MB) = 590 MB. Cost: Easily fits on a 1GB instance. Cost: ~$15/month. Highly cost-effective. Scenario B: Global Gaming Leaderboard (Sorted Sets)\nData: 10,000,000 players. Payload: Short Player ID (20 bytes) + Score (8 bytes). Math: Raw Data (280 MB) + Skip-List Overhead (~1.2 GB) = ~1.48 GB. Takeaway: Sorted sets require massive pointer overhead to maintain their $O(\\log(N))$ sorting speed. You would need a 2GB or 4GB server (~$35-$70/month). Pro Tip: Always leave a 25% memory buffer on your server for background operations (like BGSAVE for disk snapshots) to prevent OOM crashes during peak loads.\n8. Scaling Redis Memory # Eventually, you will hit the limits of a single machine. You have two paths:\nVertical Scaling (Scaling Up without Downtime) # If you need to move from a 4GB server to an 8GB server, you don\u0026rsquo;t have to take your system offline.\nProvision the new 8GB server. Configure it as a Replica of your existing 4GB Master server. Wait for the data to seamlessly sync in the background. Execute a rapid failover: Break the replication bond (REPLICAOF NO ONE) to make the 8GB server the new Master, and update your application\u0026rsquo;s DNS/Connection string to point to the new box. (Managed services like AWS ElastiCache do this for you with a click of a button). Horizontal Scaling (Redis Cluster) # When a single machine simply isn\u0026rsquo;t big enough (e.g., you need 500GB of RAM), you transition to Redis Cluster. Redis Cluster automatically shards (partitions) your data across multiple Redis nodes. It uses a concept called \u0026ldquo;Hash Slots\u0026rdquo; (there are exactly 16,384 of them). When you save a key, Redis hashes the key name to assign it to a slot, and distributes those slots evenly across your cluster of machines. This allows you to scale reads, writes, and memory capacity infinitely.\n","date":"29 March 2026","externalUrl":null,"permalink":"/blogs/redis/","section":"Blogs","summary":"","title":"Intro to Redis: Architecture, Patterns, and Scaling in Production","type":"blogs"},{"content":"If you’ve ever forcefully deleted a repository and re-cloned it just to fix a merge conflict, you are not alone. Git is notorious for its steep learning curve. The problem isn\u0026rsquo;t that Git\u0026rsquo;s commands are too complex; it\u0026rsquo;s that most developers learn the commands without understanding the underlying data model.\nOnce you stop treating Git like a black box of magic commands and start seeing it as a beautifully structured graph, everything clicks. Let\u0026rsquo;s break it down.\nWhat is Git? # At its core, Git is a Distributed Version Control System (DVCS). It allows multiple people to work on the same codebase simultaneously, tracks every modification, and lets you travel back in time to any previous state of your project.\nUnlike older systems that stored a base file and a list of changes (deltas), Git thinks of your data more like a series of snapshots. Every time you commit, Git essentially takes a picture of what all your files look like at that exact moment and stores a reference to that snapshot.\nThe Git Architecture \u0026amp; Mental Model # To truly master Git, you need to understand two key architectural concepts: the Three Trees and the Graph.\n1. The Three Trees (Local Architecture) # When you work on a repository locally, your files move through three distinct phases:\nThe Working Directory: This is your current workbench. It contains the actual files you are editing right now. The Staging Area (The Index): Think of this as a loading dock. You purposefully move modified files here (git add) to prepare them for your next commit. The Local Repository: Once you commit (git commit), Git takes everything on the loading dock and stores it permanently as a snapshot in your history. 2. The Mental Model: The Directed Acyclic Graph (DAG) # Git history is not a simple list; it\u0026rsquo;s a tree-like graph.\nCommits are nodes (circles) in the graph. Each commit points to its parent(s). Branches are not physical containers of commits. They are simply lightweight, movable pointers stuck to specific commits. HEAD is a special pointer that tells Git exactly where you are sitting in the graph right now. Usually, HEAD points to a branch name. Base State Graph: You are working on feature, which is three commits ahead of main.\n(A) --- (B) --- (C) \u0026lt;-- [main] \\ (D) --- (E) --- (F) \u0026lt;-- [feature] \u0026lt;-- [HEAD] Basic Commands: The Merge Graph # Let’s look at standard branching and merging. When you are on main and run git merge feature, Git creates a new merge commit (M) that has two parents, tying the histories together.\nInitial State:\n(A) --- (B) --- (C) \u0026lt;-- [main] \u0026lt;-- [HEAD] \\ (D) --- (E) \u0026lt;-- [feature] Action: git merge feature\nResulting Graph: A new commit (M) is created, combining changes. The [main] pointer moves to M.\n(A) --- (B) --- (C) --- (M) \u0026lt;-- [main] \u0026lt;-- [HEAD] \\ / (D) ------- (E) \u0026lt;-- [feature] Undoing Mistakes: Git Reset vs. Git Revert # Both commands undo changes, but they affect the graph in entirely different ways.\ngit reset (Rewriting History) # Reset is used for local, unpushed changes. It picks up the current branch pointer and moves it backward in time.\nInitial State: You realize commit (C) is broken and you want to erase it.\n(A) --- (B) --- (C) \u0026lt;-- [main] \u0026lt;-- [HEAD] Action: git reset --hard B (Moves pointer to B and clears working directory).\nResulting Graph: The [main] pointer moves back to B. Commit (C) is left behind. It is now orphaned and will eventually be deleted by Git\u0026rsquo;s garbage collection. History has been rewritten.\n(A) --- (B) \u0026lt;-- [main] \u0026lt;-- [HEAD] [ (C) orphaned ] git revert (Adding to History) # Revert is used for shared remote branches. It does not erase commits; instead, it figures out the exact opposite of a specific commit and adds a new commit performing that undo action.\nInitial State: Commit (C) is broken, but you already pushed it to the remote. You can\u0026rsquo;t use reset.\n(A) --- (B) --- (C) \u0026lt;-- [main] \u0026lt;-- [HEAD] Action: git revert C\nResulting Graph: Git creates a new commit (C_Opposite) which cancels out the changes made in (C). The history flows forward cleanly.\n(A) --- (B) --- (C) --- (C_Opposite) \u0026lt;-- [main] \u0026lt;-- [HEAD] Detached HEAD and Relative Refs # The Detached HEAD Graph # Normally, HEAD points to a branch name (e.g., [main]). If you check out a specific commit hash (e.g., git checkout B), HEAD detaches from the branch and points directly to the commit.\nInitial State:\n(A) --- (B) --- (C) \u0026lt;-- [main] \u0026lt;-- [HEAD] Action: git checkout B (detached HEAD state). If you make a new commit (X) here, it branches off into a void.\nResulting Graph:\n[HEAD] | (X) \u0026lt;-- (New orphaned commit) / (A) --- (B) --- (C) \u0026lt;-- [main] (If you switch back to main now, commit X is lost. To save it, you must create a branch right there: git switch -c recovery).\nRelative References Graph (~ and ^) # Popularized by Learn Git Branching, relative references make navigating the graph easy. Assume this is our graph with HEAD at commit (D):\n(A) --- (B) --- (C) \u0026lt;-- [main] \\ (D) \u0026lt;-- [feature] \u0026lt;-- [HEAD] HEAD~1 (Parent) = (B) HEAD~2 (Grandparent) = (A) If (D) were a merge commit with parents (B) and (C):\nHEAD^1 (First parent, straight back) = (B) HEAD^2 (Second parent, the branch merged in) = (C) The Ultimate Safety Net: Git Log and Reflog # git log: Shows the official history of the current branch, tracing parent nodes backward. git reflog (Reference Log): The \u0026ldquo;time travel\u0026rdquo; insurance policy. git log only shows active commits. git reflog lists every single movement of the HEAD pointer on your local machine, whether you switched branches, reset, or committed. Even if you --hard reset a branch and seemingly lose all your work, you can find the lost commit hash in the reflog and reset right back to it. Advanced Superpowers: Cherry-Pick and Interactive Rebase # Git Cherry-Pick Graph # Cherry-pick acts like a cut-and-paste tool. It allows you to grab a specific commit from one branch and apply it to another.\nInitial State: You are on main and want the specific hotfix made in commit (E) on the feature branch, without merging all the other unfinished work.\n(A) --- (B) --- (C) \u0026lt;-- [main] \u0026lt;-- [HEAD] \\ (D) --- (E) --- (F) \u0026lt;-- [feature] Action: git cherry-pick E\nResulting Graph: Git copies the changes from (E) and creates a brand new commit (E\u0026rsquo;) on main. It has a different hash because it has a different parent.\n(A) --- (B) --- (C) --- (E\u0026#39;) \u0026lt;-- [main] \u0026lt;-- [HEAD] \\ (D) --- (E) --- (F) \u0026lt;-- [feature] Interactive Rebase Graph (Squashing) # Interactive rebase (git rebase -i) allows you to pause time and rewrite a series of commits before they are finalized. It is commonly used to \u0026ldquo;squash\u0026rdquo; multiple messy commits into one.\nInitial State: You have finished a feature, but you made 3 messy commits (wip, fix typo, actually blue).\n(A) --- (B) (main) \\ (C) --- (D) --- (E) \u0026lt;-- [feature] \u0026lt;-- [HEAD] ^ ^ ^ \u0026#34;wip\u0026#34; \u0026#34;typo\u0026#34; \u0026#34;done\u0026#34; Action: git rebase -i main (We choose to squash D and E into C).\nResulting Graph: Git melts commits C, D, and E down and creates a single, clean commit (F).\n(A) --- (B) (main) \\ (F) \u0026lt;-- [feature] \u0026lt;-- [HEAD] ^ \u0026#34;feat: add button\u0026#34; Common use cases for rebasing # 1. Keep your feature branch up to date with main # Situation # You started a feature branch a few days ago:\nmain: A---B---C \\ feature: D---E Meanwhile, main moved forward:\nmain: A---B---C---F---G \\ feature: D---E What you do # git checkout feature git rebase main Result # main: A---B---C---F---G \\ feature: D\u0026#39;---E\u0026#39; Why rebase here? # Keeps history linear Avoids a messy merge commit Makes it look like you started from the latest code Use this when:\nYou\u0026rsquo;re working alone on the branch You want clean history before merging 2. Clean up messy commits before merging (interactive rebase) # Situation # Your commits look like this:\nD - \u0026#34;fix\u0026#34; E - \u0026#34;oops\u0026#34; F - \u0026#34;final fix\u0026#34; What you do # git rebase -i main Then squash:\npick D squash E squash F Result # D\u0026#39; - \u0026#34;Add login feature\u0026#34; Why rebase here? # Turns messy work into clean, meaningful commits Makes code review easier Use this when:\nYou\u0026rsquo;re about to open a PR You want professional-looking history Why rebase here? # Keeps dependency chain intact\nAvoids weird merge commits between features Use this when:\nYou have layered branches (very common in big features)\n3. Before merging into main (linear history teams) # Some teams prefer:\nRebase + fast-forward merge No merge commits at all Workflow # git checkout feature git rebase main git checkout main git merge feature # fast-forward Result # main: A---B---C---D\u0026#39;---E\u0026#39; Git is powerful, and with the right mental model—seeing it as a graph of snapshots governed by moving pointers—you can navigate any version control crisis with confidence.\nReference # The Missing Semester of Your CS Education Learngitbranching.js.org ","date":"28 March 2026","externalUrl":null,"permalink":"/blogs/git_basics/","section":"Blogs","summary":"","title":"Git: Building the Right Mental Model","type":"blogs"},{"content":" In the classical era of system administration, provisioning a server meant logging into a web console, clicking dozens of buttons, and hoping you remembered every setting. If you needed ten identical servers, you repeated that process ten times. This approach was slow, prone to human error, and impossible to scale.\nThen came Infrastructure as Code (IaC), and Terraform emerged as the industry standard.\n1. What’s Terraform and Why Terraform? # Terraform is an open-source tool created by HashiCorp. It is an Infrastructure as Code (IaC) tool that allows you to define, provision, and manage both cloud and on-premise resources using human-readable configuration files.\nWhy is Terraform the Go-To Tool? # The primary reason is that Terraform is Declarative. You don\u0026rsquo;t write scripts telling the cloud how to build a server (like standard Bash or Python scripts). Instead, you write a configuration defining what the final infrastructure must look like. Terraform figures out the execution steps to reach that desired state.\nOther key benefits include:\nCloud-Agnostic: It works seamlessly with AWS, Azure, Google Cloud, Kubernetes, VMware, and hundreds of other providers using the same workflow. Immutable Infrastructure: It favors replacing resources over modifying them, reducing configuration drift. Idempotence: You can run the same Terraform plan multiple times, and it will only apply the changes necessary to reach the final state. Running an identical plan a second time changes nothing. 2. Terraform Components and Architecture # To understand Terraform, you have to understand its internal anatomy.\nThe Engine and Core # Terraform’s Core is the main executable binary (written in Go). It\u0026rsquo;s responsible for the overall logic. It parses your configuration files, manages the state, and creates the Resource Graph—a dependency map that determines the correct order to create or destroy resources (e.g., ensuring the Network is built before the Virtual Machine).\nHow Providers Work # Terraform Core doesn\u0026rsquo;t actually know how to talk to AWS or Azure API. It uses Providers, which are plugins. The core communicates with these plugins via RPC (Remote Procedure Calls).\nWhen you run an apply, Terraform sends generic instructions to the AWS Provider, which translates them into specific AWS API calls (like ec2:RunInstances). This plugin architecture is what allows Terraform to manage almost any service with an API.\nState Management: The Source of Truth # The most critical component is the State File ($terraform.tfstate$). This JSON file acts as Terraform\u0026rsquo;s memory. It maps the resources defined in your code to the actual, real-world resources currently existing in your cloud account.\nWhen you run a plan, Terraform compares your code against this state file to determine what needs to be changed. Protect this file: if you lose it, Terraform will lose track of your managed infrastructure.\n3. Terraform Workflow # The classic Terraform workflow is a simple, iterative three-step process:\nWrite: You write your infrastructure definitions in .tf files. Plan (terraform plan): The engine compares your code to the state file and generates an execution plan. It tells you exactly which resources will be + Created, ~ Updated, or - Destroyed. This is your dry run. Apply (terraform apply): Upon confirmation, Terraform calls the required providers to execute the plan and update the state file. 4. Every Code Component with Examples # Let’s look at the foundational blocks of a Terraform project, using AWS for our examples.\nThe Providers Block # Configures the plugin.\n# providers.tf provider \u0026#34;aws\u0026#34; { region = \u0026#34;us-east-1\u0026#34; } Variables # Allows you to inject dynamic inputs and avoid hard-coding.\n# variables.tf variable \u0026#34;instance_type\u0026#34; { type = string default = \u0026#34;t2.micro\u0026#34; description = \u0026#34;The size of the server\u0026#34; } Locals # Internal, private variables used for cleanup or calculations. The user cannot override these.\n# main.tf (locals block) locals { # Logic to enforce standardized naming conventions project_prefix = \u0026#34;ecommerce-prod\u0026#34; server_name = \u0026#34;${local.project_prefix}-webserver\u0026#34; } Data Sources # Queries existing infrastructure or dynamic information from the cloud provider (Read-Only).\n# main.tf (data block) data \u0026#34;aws_ami\u0026#34; \u0026#34;latest_ubuntu\u0026#34; { most_recent = true owners = [\u0026#34;099720109477\u0026#34;] # Canonical/Ubuntu filter { name = \u0026#34;name\u0026#34; values = [\u0026#34;ubuntu/images/hvm-ssd/ubuntu-focal-20.04-amd64-server-*\u0026#34;] } } Resources # The most important part—this is what you are actually building.\n# main.tf (resource block) resource \u0026#34;aws_instance\u0026#34; \u0026#34;web\u0026#34; { # Referencing the DATA SOURCE result ami = data.aws_ami.latest_ubuntu.id # Referencing the VARIABLE instance_type = var.instance_type # Referencing the LOCAL tags = { Name = local.server_name } } Outputs # Prints specific information to your terminal after an apply.\n# outputs.tf output \u0026#34;web_server_public_ip\u0026#34; { value = aws_instance.web.public_ip description = \u0026#34;Connect to the server at this IP\u0026#34; } More about Terraform providers # In Terraform, a provider is a plugin that interacts with an external system’s API to manage resources.\nIt implements CRUD operations:\nCreate Read Update Delete Types of Terraform Providers # Terraform providers are not officially classified this way, but in practice they fall into four functional categories:\n1. Infrastructure Providers # Definition # Providers that manage core infrastructure resources such as compute, networking, and storage.\nExamples # Amazon Web Services via AWS provider Microsoft Azure Google Cloud Platform Docker Example # resource \u0026#34;aws_instance\u0026#34; \u0026#34;web\u0026#34; { ami = \u0026#34;ami-123\u0026#34; instance_type = \u0026#34;t2.micro\u0026#34; } Characteristics # Directly create infrastructure Strong state mapping High reliability Widely used Use cases # VMs, VPCs, load balancers Storage systems Base infrastructure setup 2 Platform Providers # Definition # Providers that manage platforms which themselves orchestrate infrastructure or workloads.\nExamples # Kubernetes OpenShift Example # resource \u0026#34;kubernetes_deployment\u0026#34; \u0026#34;app\u0026#34; { metadata { name = \u0026#34;nginx\u0026#34; } } Characteristics # Operate on top of infrastructure Interact with platform APIs Manage workloads, not raw infra Use cases # Deploying applications Managing cluster resources Service definitions 3 Orchestration / Packaging Providers # Definition # Providers that manage deployment tools or packaging systems, not raw infrastructure.\nExamples # Helm (conceptually similar tools: ArgoCD, though not always used via Terraform) Example # resource \u0026#34;helm_release\u0026#34; \u0026#34;nginx\u0026#34; { name = \u0026#34;nginx\u0026#34; chart = \u0026#34;nginx\u0026#34; } Characteristics # Wrap another system Often depend on another provider (e.g., Kubernetes) Manage packaged applications Use cases # Installing applications via charts Managing versions of deployed apps 4 Application / Service Providers # Definition # Providers that manage configuration inside SaaS tools or applications.\nExamples # Grafana Datadog GitHub Example # resource \u0026#34;grafana_dashboard\u0026#34; \u0026#34;example\u0026#34; { config_json = file(\u0026#34;dashboard.json\u0026#34;) } Characteristics # Manage logical configuration, not infrastructure API-driven Often sensitive to drift (manual UI changes) Use cases # Dashboards, alerts Repositories, permissions Monitoring configuration 5 Utility / Bridge Providers # Definition # Providers that bridge gaps or provide utility functionality, often mimicking CLI behavior.\nExamples # Community kubectl provider (wraps kubectl behavior) Random provider (generates random values) Null provider (executes scripts) Example # resource \u0026#34;kubectl_manifest\u0026#34; \u0026#34;example\u0026#34; { yaml_body = file(\u0026#34;manifest.yaml\u0026#34;) } Characteristics # Often community-built May be imperative in nature Weaker state guarantees Use cases # Applying raw YAML Running scripts Handling unsupported resources Provider Ownership Types # Another important classification is who maintains the provider:\nOfficial Providers # Maintained by HashiCorp\nExamples:\nhashicorp/aws hashicorp/kubernetes hashicorp/helm Vendor Providers # Maintained by the platform company\nExamples:\ngrafana/grafana datadog/datadog Community Providers # Maintained by individuals or small groups\nExamples:\nkreuzwerker/docker gavinbunney/kubectl Key Differences Across Types # Type Manages State Complexity Example Infrastructure Physical/cloud resources Low AWS Platform Workloads/platform objects Medium Kubernetes Orchestration Packaged deployments Medium–High Helm Application App-level config Medium Grafana Utility Helpers/bridges High kubectl By mastering these fundamental building blocks, you have the key to defining and managing entire data centers in simple, version-controlled text files. Happy provisioning!\n","date":"27 March 2026","externalUrl":null,"permalink":"/blogs/terraform/","section":"Blogs","summary":"","title":"Terraform: Your Guide to Infrastructure as Code","type":"blogs"},{"content":"Modern applications rely heavily on databases to store and manage critical data. But with multiple users, concurrent operations, and potential system failures, how can we ensure that data remains accurate, consistent, and reliable?\nThis is where transactions and the ACID properties come in. A transaction allows a set of database operations — reads, writes, updates, and deletes — to be executed as a single logical unit. Either all operations succeed (commit) or none take effect (rollback), providing a safety net against partial updates, crashes, and concurrency issues.\nBy enforcing Atomicity, Consistency, Isolation, and Durability, databases guarantee predictable behavior even in complex, concurrent environments. At the same time, developers must understand the trade-offs between strict data guarantees and system performance, especially when scaling or replicating data.\nIn this blog, we’ll explore:\nHow transactions work The meaning and examples of ACID properties Isolation levels and their impact on concurrent transactions Practical guidance for choosing the right level of consistency for your applications By the end, you’ll have a clear understanding of how databases maintain integrity, how to avoid anomalies, and how to select the best strategy for your application’s needs.\nWhat Are Transactions? # A transaction is a sequence of one or more read and write operations grouped into a single unit of work. It either commits (successfully applies all changes) or aborts/rolls back (undoes all changes as if nothing happened).\nExample # BEGIN TRANSACTION; UPDATE accounts SET balance = balance - 500 WHERE id = 1; UPDATE accounts SET balance = balance + 500 WHERE id = 2; COMMIT; If the system crashes after the first update but before the second, the transaction is rolled back — ensuring that no account loses money.\nTransactions aren’t laws of nature — they’re a design choice made to ensure data integrity and ease of development. However, as modern NoSQL systems emerged (e.g., MongoDB, Cassandra), many relaxed or dropped these guarantees to achieve scalability, replication, and partitioning by default. In many cases, transactions were either absent or weakened (offering only per-document atomicity).\nSingle-Object Transactions # Even when modifying a single data object, atomicity and isolation remain critical.\nImagine writing a 20 KB JSON document to a database:\nIf the network disconnects after sending 10 KB, will the database store a broken JSON fragment? If the power fails midway while overwriting the old value, could the old and new data be spliced together? If another client reads the document during the write, will it see a partially written version? In databases like PostgreSQL or MySQL (InnoDB), the answers are:\nNo partial data is ever stored or visible. The operation is atomic — you see either the old or the new value, never a mix. Thus, every single write behaves as an all-or-nothing operation, even for a single object.\nMulti-Object Transactions # When an application modifies multiple rows, tables, or documents, it uses multi-object transactions.\nEverything between a BEGIN TRANSACTION and COMMIT statement is part of a single atomic operation:\nBEGIN TRANSACTION; INSERT INTO orders VALUES (101, \u0026#39;MacBook Pro\u0026#39;, 2499); UPDATE inventory SET quantity = quantity - 1 WHERE product_id = \u0026#39;MBP\u0026#39;; COMMIT; If any statement in between fails — for example, if the inventory update violates a constraint — the transaction rolls back, and the new order is not saved.\nIn SQL systems, a multi-object transaction is scoped to the client’s TCP connection with the database. All statements issued through that connection between BEGIN TRANSACTION and COMMIT are treated as part of one logical transaction.\nBy using transactions, applications gain:\nProperty Description Atomicity Ensures that operations are all-or-nothing. Consistency Guarantees that data remains valid before and after the transaction. Isolation Prevents concurrent transactions from interfering with one another. Durability Ensures committed data survives failures. These principles — known as ACID properties — form the foundation of reliable transactional systems.\nACID Properties in Transactions # Transactions in databases are governed by four key properties — Atomicity, Consistency, Isolation, and Durability, collectively known as ACID. These properties ensure that database operations remain reliable even in the face of system crashes, concurrent users, or unexpected errors.\nA — Atomicity (Abortability) # Definition: Atomicity ensures that a transaction is all or nothing — either every operation in the transaction succeeds, or none of them do. If an error occurs midway, the database automatically aborts the transaction and rolls back any partial changes.\nThis property eliminates the problem of partial failure, where some changes are saved but others are not.\nExample # Consider a fund transfer between two accounts:\nBEGIN TRANSACTION; UPDATE accounts SET balance = balance - 500 WHERE id = 1; UPDATE accounts SET balance = balance + 500 WHERE id = 2; COMMIT; If the second update fails due to a network issue or constraint error, the first update (debit from account 1) is rolled back.\nResult # Step Action Outcome Step 1 Withdraw 500 from Account 1 Success Step 2 Deposit 500 to Account 2 Fails Result Transaction rolled back Account 1 balance restored, Account 2 unchanged The database guarantees that no money disappears or appears out of thin air — maintaining atomicity.\nC — Consistency # Definition: Consistency ensures that a transaction moves the database from one valid state to another. It enforces integrity constraints such as foreign keys, unique indexes, or business rules.\nWhile the database enforces physical consistency (like constraints and data types), the application logic often enforces logical consistency (like valid business rules).\nExample # Suppose there’s a constraint that the total sum of balances across all accounts must remain constant:\nALTER TABLE accounts ADD CONSTRAINT positive_balance CHECK (balance \u0026gt;= 0); Now, if a transaction tries to withdraw more money than the account holds:\nBEGIN TRANSACTION; UPDATE accounts SET balance = balance - 2000 WHERE id = 1; -- Account has only 1000 COMMIT; The CHECK constraint fails, and the transaction is aborted.\nResult # Transaction Step Action Status Debit exceeds balance Violates constraint Transaction rolled back Final state Database unchanged Consistent Even if the programmer forgets to validate input, the database won’t commit inconsistent data.\nI — Isolation # Definition: Isolation ensures that concurrent transactions do not interfere with one another. Each transaction behaves as if it were the only one executing on the system, even when many run at the same time.\nWithout isolation, concurrent operations can cause problems such as:\nDirty reads – Reading uncommitted data from another transaction. Non-repeatable reads – Seeing different results for the same query in one transaction. Phantom reads – Rows appearing or disappearing during a transaction due to concurrent inserts/deletes. Example # Two users simultaneously try to update the same account:\n-- Transaction A BEGIN TRANSACTION; UPDATE accounts SET balance = balance - 100 WHERE id = 1; -- Transaction B BEGIN TRANSACTION; UPDATE accounts SET balance = balance - 50 WHERE id = 1; If isolation is weak (e.g., READ UNCOMMITTED), both might read the same initial balance, resulting in incorrect final values.\nWith proper isolation (e.g., REPEATABLE READ or SERIALIZABLE), the database serializes these updates — one completes before the other starts.\nResult # Isolation Level Behavior Possible Issue READ UNCOMMITTED Transactions can see uncommitted data Dirty read READ COMMITTED Reads only committed data Non-repeatable reads REPEATABLE READ Prevents dirty and non-repeatable reads Phantom reads possible SERIALIZABLE Transactions fully isolated Slowest but safest We’ll explore these isolation levels in depth in the next section.\nD — Durability # Definition: Durability ensures that once a transaction is committed, its results are permanently saved, even if the system crashes immediately afterward.\nDatabases achieve durability by writing data to non-volatile storage (like disk logs or SSDs) before confirming the commit to the client.\nDurability ensures that once a transaction is committed, its effects are permanently recorded on stable storage or safely replicated across nodes. Databases achieve this using techniques like write-ahead logging, checkpointing, and replication. Even after crashes or power failures, committed data is recoverable and never lost.\nExample # BEGIN TRANSACTION; UPDATE inventory SET quantity = quantity - 1 WHERE product_id = \u0026#39;MBP\u0026#39;; COMMIT; Once the commit succeeds, the update is stored in the write-ahead log (WAL). If the server crashes right after, the database recovers using the log and replays the committed transaction during restart.\nResult # Step Event Outcome Commit acknowledged WAL written to disk Safe Server crash Database restarts Transaction recovered Final state Data persists Durable Isolation Level: READ COMMITTED # Explanation # The READ COMMITTED isolation level guarantees two key properties:\nNo Dirty Reads: A transaction never reads data that another transaction has written but not yet committed. No Dirty Writes: A transaction only overwrites data that has been committed, never uncommitted updates. This is the default isolation level in many databases such as PostgreSQL and Oracle. Each statement within a transaction sees only data that was committed before that statement began.\nImplementation # Databases typically implement READ COMMITTED using row-level locks and versioning:\nFor each object (row) being written, the database stores both:\nThe old committed value (visible to other transactions). The new uncommitted value (visible only to the transaction that holds the write lock). When a transaction commits, its new values replace the old committed ones.\nOther transactions reading the same data will always see the most recently committed version, ensuring that no uncommitted data is exposed.\nExample # Let’s simulate two concurrent transactions operating on the same account balance.\n-- Transaction T1 BEGIN TRANSACTION; SELECT balance FROM accounts WHERE id = 1; -- Reads balance = 1000 -- (Some delay before next step) -- Transaction T2 BEGIN TRANSACTION; UPDATE accounts SET balance = balance - 200 WHERE id = 1; COMMIT; -- Transaction T1 continues SELECT balance FROM accounts WHERE id = 1; -- Reads balance = 800 COMMIT; Result:\nT1’s first read returns 1000 (old committed value). T2 commits and changes the balance to 800. When T1 reads again, it sees 800 — the new committed value. This behavior ensures no dirty reads, but T1 sees different values within the same transaction — a phenomenon known as non-repeatable read.\nUnhandled Cases (Anomalies under READ COMMITTED) # Although READ COMMITTED prevents dirty reads and writes, it still allows several anomalies:\n1. Read Skew / Non-Repeatable Read # Two concurrent transactions start at the same time, read the same value, but one commits and modifies the value before the other finishes. When the first transaction re-reads the value, it sees a new committed version, producing inconsistent results.\nImpact:\nCauses inaccurate results in backups, analytics, or integrity checks where consistent snapshots are needed. Example: While generating a financial report, the total balance changes midway through the read. 2. Lost Updates # Two transactions read the same committed value and both attempt to update it:\n-- T1 BEGIN; SELECT counter FROM stats WHERE id = 1; -- Reads 5 UPDATE stats SET counter = 6 WHERE id = 1; COMMIT; -- T2 BEGIN; SELECT counter FROM stats WHERE id = 1; -- Reads 5 UPDATE stats SET counter = 7 WHERE id = 1; COMMIT; Result: Both read the same initial value (5), but the second commit overwrites the first — resulting in one lost update.\n3. Write Skew # A subtler anomaly where two transactions update different objects based on overlapping reads.\nScenario: Two doctors, Alice and Bob, are on call. At least one doctor must remain on call at all times.\nDoctor on_call Alice TRUE Bob TRUE Now both doctors attempt to go off call concurrently.\n-- Transaction T1 (Alice) BEGIN; SELECT COUNT(*) FROM doctors WHERE on_call = TRUE; -- Returns 2 UPDATE doctors SET on_call = FALSE WHERE name = \u0026#39;Alice\u0026#39;; COMMIT; -- Transaction T2 (Bob) BEGIN; SELECT COUNT(*) FROM doctors WHERE on_call = TRUE; -- Returns 2 UPDATE doctors SET on_call = FALSE WHERE name = \u0026#39;Bob\u0026#39;; COMMIT; Expected Behavior: At least one doctor should stay on call.\nActual Behavior: Both transactions read the same committed state (2 doctors on call) and proceed independently. After both commits, no one is on call — a logical inconsistency known as write skew.\nProperty Ensures Fails To Prevent Use Cases READ COMMITTED No dirty reads, No dirty writes Non-repeatable reads, Lost updates, Write skew Most OLTP systems where perfect consistency across reads isn’t required Repeatable Read Isolation Level # Explanation # Repeatable Read (Snapshot Isolation) ensures that all reads within a transaction see a consistent snapshot of the database as it existed at the time the transaction started. Even if other transactions commit changes after it begins, the running transaction continues to see only its own snapshot view.\nThis level solves problems like read skew (where different reads within the same transaction return inconsistent results) by guaranteeing that data read once will not change if read again later in the same transaction.\nImplementation # Snapshot isolation is most commonly implemented using Multi-Version Concurrency Control (MVCC). Key implementation principles include:\nReaders never block writers, and writers never block readers. Each transaction sees data that was committed before it started, ignoring uncommitted or later transactions. Write locks are used to prevent dirty writes, while reads require no locks. The database maintains multiple versions of an object to serve transactions starting at different times. Updates are often implemented as a delete + create operation internally to maintain version history. Visibility Rules # When a transaction starts, the database determines which committed versions are visible:\nIgnore writes from transactions that were in-progress when the current transaction started. Ignore writes from aborted transactions. Ignore writes from transactions that started after the current transaction. All other committed writes are visible. Example # Scenario: Two analysts run transactions on a salary table.\nT1: BEGIN TRANSACTION T1: SELECT SUM(salary) FROM employees; -- returns 200K T2: BEGIN TRANSACTION T2: UPDATE employees SET salary = salary + 5000 WHERE id=1; T2: COMMIT T1: SELECT SUM(salary) FROM employees; -- still returns 200K (unchanged) T1: COMMIT Result: Even though T2 committed an update, T1 continues to see the consistent snapshot of data (as of its start). This prevents read skew and ensures repeatable reads.\nUnhandled Cases # Despite solving several read anomalies, Repeatable Read does not address all concurrency conflicts — particularly those involving writes.\n1. Lost Updates # Occurs when two concurrent transactions perform read-modify-write operations on the same data. Example:\nT1: SELECT counter FROM metrics WHERE key=\u0026#39;foo\u0026#39;; -- returns 10 T2: SELECT counter FROM metrics WHERE key=\u0026#39;foo\u0026#39;; -- returns 10 T1: UPDATE metrics SET counter=11 WHERE key=\u0026#39;foo\u0026#39;; T2: UPDATE metrics SET counter=11 WHERE key=\u0026#39;foo\u0026#39;; Result: One update overwrites the other, losing a modification.\nSolutions:\nAtomic write operations:\nUPDATE metrics SET counter = counter + 1 WHERE key = \u0026#39;foo\u0026#39;; Explicit locking before read-modify-write:\nSELECT * FROM metrics WHERE key=\u0026#39;foo\u0026#39; FOR UPDATE; 2. Phantoms (Write Skew) # A phantom read occurs when one transaction’s write changes the result set of another’s query. For example, in a hospital scheduling system:\nT1: SELECT COUNT(*) FROM doctors WHERE on_call = true; T2: SELECT COUNT(*) FROM doctors WHERE on_call = true; -- Both see 1 doctor on call. T1: UPDATE doctors SET on_call=false WHERE id=1; T2: UPDATE doctors SET on_call=false WHERE id=2; -- Both commit -\u0026gt; 0 doctors on call. Result: System invariant (at least one doctor on call) is violated.\nSolution: Use materializing conflicts — introduce a lock on a concrete set of rows or use higher isolation levels like Serializable to detect and abort conflicting transactions.\n3. Write Skew in Derived Constraints # Even though transactions don’t update the same row, they indirectly violate constraints that depend on multiple rows (e.g., scheduling, inventory allocation).\nMitigation: Convert logical constraints into explicit locking or enforce them through check constraints and serializable isolation.\nSerializable Isolation Level # Explanation # Serializable isolation is the strongest isolation level in database systems. It guarantees that even though transactions run concurrently, the end result is equivalent to executing them one after another in a serial order.\nIn other words, no set of concurrent transactions can produce an effect that could not occur if those same transactions were executed sequentially.\nThis level eliminates all read and write anomalies — including dirty reads, non-repeatable reads, phantoms, and write skew — ensuring full consistency.\nImplementations # There are three primary strategies for achieving serializability in modern databases, each with different trade-offs between performance and isolation strength.\n1. Literal Serial Execution # Approach: Transactions are executed one at a time in a single thread, guaranteeing serializability by design.\nFeasibility factors:\nWith modern large memory systems, most OLTP datasets can fit entirely in RAM. OLTP transactions are typically short and touch a small subset of records. Therefore, running them serially eliminates concurrency issues without noticeable slowdown. Examples:\nImplemented by systems like VoltDB, H-Store, Redis, and Datomic, where transactions are short and deterministic. Result: Since only one transaction executes at a time, there are no conflicts, no locks, and no rollbacks due to concurrency — achieving perfect serializability naturally.\n2. Two-Phase Locking (2PL) # Approach: Two-phase locking enforces strict lock acquisition and release rules to ensure no conflicting transactions proceed simultaneously.\nMechanism: # Each database object (row, record, etc.) is associated with a lock:\nShared lock (S): multiple transactions can hold it for reading. Exclusive lock (X): only one transaction can hold it for writing. Locking Rules: # To read, acquire a shared lock.\nMultiple transactions can share this lock. To write, acquire an exclusive lock.\nBlocks both readers and writers until released. Locks are held until the transaction completes (commit or abort).\nHence, “two phases”:\nGrowing phase: acquiring locks. Shrinking phase: releasing all locks at the end. Example: # T1: READ account_balance (shared lock) T2: tries WRITE account_balance (exclusive lock) → must wait T1: COMMIT → releases lock T2: proceeds with WRITE Result: No conflicting reads or writes can occur simultaneously — transactions are effectively serialized by lock dependencies.\nLimitations: # Reduced concurrency: readers block writers and vice versa. Deadlocks may occur if transactions wait for each other’s locks. Phantom Problem: occurs when a transaction reads a range of rows that another transaction later inserts into. Solution: Use predicate locks or index-range locks to prevent phantom writes by locking the range condition rather than individual rows.\n3. Serializable Snapshot Isolation (SSI) # Approach: A modern, optimistic concurrency control technique combining snapshot isolation (MVCC) with serializability checks.\nUnlike two-phase locking, SSI allows transactions to proceed without blocking and detects conflicts only at commit time.\nKey Concepts: # Transactions execute on consistent snapshots (from MVCC).\nAt commit, the system checks for serialization anomalies such as:\nStale reads: a transaction read outdated data that was later modified. Write-read conflicts: one transaction writes data that another previously read. If such conflicts are found, one of the transactions is aborted and retried.\nExample: # T1: Reads product inventory (50) T2: Reads product inventory (50) T1: Decreases inventory by 10 → writes 40 T2: Decreases inventory by 10 → writes 40 -- Conflict detected at commit time → one transaction aborted Result: Although transactions appear to execute concurrently, the system ensures only one final serializable order of execution is accepted.\nAdvantages: # High concurrency like snapshot isolation. Guarantees full serializability without long lock waits. Disadvantages: # Transactions may be aborted and retried frequently under high contention. Implementations: # Found in PostgreSQL (Serializable Snapshot Isolation mode), FoundationDB, and CockroachDB. Approach Concurrency Blocking Common In Notes Literal Serial Execution None None VoltDB, Redis Simplest, works for in-memory OLTP Two-Phase Locking (2PL) Low High (readers/writers block) MySQL (InnoDB), Oracle Classical, robust but slower Serializable Snapshot Isolation (SSI) High None (optimistic) PostgreSQL, CockroachDB Modern, best balance between safety \u0026amp; concurrency Outcome # Serializable isolation is the gold standard for correctness, ensuring complete transactional integrity. However, it comes with trade-offs — reduced throughput (in 2PL) or occasional rollbacks (in SSI). Choosing the right implementation depends on your workload:\nShort, high-speed OLTP: literal serial execution. Complex, high-contention workloads: 2PL. Mixed read/write workloads: SSI for balance between speed and correctness. Comparison of Isolation Levels # Isolation Level Prevents Dirty Reads Prevents Non-Repeatable Reads Prevents Phantoms / Write Skew Concurrency Typical Use Case READ COMMITTED Yes No No High Most OLTP systems where perfect repeatable reads are not required REPEATABLE READ Yes Yes No (Write Skew possible) Medium Systems needing repeatable reads, e.g., analytics or reporting SERIALIZABLE Yes Yes Yes Low (2PL) / Medium (SSI) Critical systems needing full consistency, e.g., banking, reservations Conclusion # ACID properties — Atomicity, Consistency, Isolation, Durability — are the foundation of reliable database systems.\nIsolation levels define how transactions interact and what anomalies are prevented, balancing consistency vs. concurrency.\nModern databases face a trade-off between scalability, replication, and partitioning versus full ACID guarantees.\nApplication use case matters: choose a datastore and isolation level based on the consistency requirements and performance needs.\nFor example:\nHigh-throughput analytics may tolerate READ COMMITTED or REPEATABLE READ. Financial or reservation systems require SERIALIZABLE isolation. Key Takeaway: Understanding ACID properties and isolation levels allows developers and architects to make informed decisions, ensuring data integrity without unnecessarily compromising performance or scalability.\nReference # Designing-data-intensive-applications - Book by Martin Kleppmann ","date":"5 October 2025","externalUrl":null,"permalink":"/blogs/transaction_isolation/","section":"Blogs","summary":"","title":"Isolation levels in Databases","type":"blogs"},{"content":"In today’s world, where applications serve millions of concurrent users, delivering fast and reliable responses is non-negotiable. That’s where load balancers come in. Think of them as invisible traffic controllers sitting between users and servers, making sure every request is handled efficiently without overloading any single server.\nWhat is Load Balancing and Why Do We Need It? # Load balancing is the process of distributing network traffic across a pool of servers (resources) that host an application.\nWithout it, one server could get overwhelmed while others sit idle, causing downtime, delays, and unhappy users.\nA load balancer ensures:\nAvailability\nSeamless maintenance without downtime Disaster recovery handling Continuous health checks to detect failed servers Scalability\nEffortlessly add or remove servers to handle traffic spikes or scale down during quiet periods In essence, load balancing keeps applications resilient, scalable, and always online.\nTypes of Load Balancing # Hardware-based Load Balancers\nDedicated appliances with high throughput, used in large enterprises. Expensive but extremely optimized. Software-based Load Balancers\nRun on commodity hardware or cloud VMs. Popular examples: HAProxy, NGINX, Envoy, Traefik. Cloud-native: AWS ELB/ALB, GCP Load Balancer, Azure Front Door. Major Categories # Application-Level (Layer 7 / L7) – HTTP, HTTPS, gRPC aware Network-Level (Layer 3/4) – Operates on IP and TCP/UDP layers Load Balancing Algorithms # Load balancers use algorithms to decide which server gets the next request.\n1. Static Methods # Round Robin – Sequentially distributes requests across servers. Weighted Round Robin – Assigns more traffic to higher-capacity servers. IP Hash – Maps client IPs consistently to the same server. 2. Dynamic Methods # Least Connections – Chooses the server with the fewest active connections. Weighted Least Connections – Accounts for server capacity in addition to connections. Least Response Time – Sends traffic to the fastest responder. Resource-Based Monitoring – Routes based on CPU/memory load. Performance Difference Between L3/L4 and L7 Load Balancers # Layer 4 (Transport Load Balancer) # Operates at TCP/UDP level. Makes decisions only on IP and port. Does not terminate connections → acts like a smart router. Minimal overhead, best for raw performance. Layer 7 (Application Load Balancer) # Operates at HTTP/HTTPS/gRPC level. Terminates client connections, decrypts TLS if needed. Parses headers (Host, Path, cookies) to make content-aware routing decisions. Opens new TCP connections to backend servers. Overhead exists because data is copied and parsed. Is there an “extra copy” of data? # Yes, at L7 → requests are buffered/parsed. No, at L4 → packets are passed through transparently. Modern proxies minimize this overhead with zero-copy system calls (sendfile(), splice()) and streaming instead of full buffering. Example: Request Flow in L7 Proxy (HAProxy / NGINX) # Client → Load Balancer\nClient sends GET /api/users. LB terminates TLS, parses headers. Load Balancer → Backend\nRoutes /api/* → backend1. Opens a new HTTP connection. Streams request body directly to backend. Backend → Load Balancer → Client\nResponse is proxied back. LB may add headers (X-Forwarded-For). Practical Implications # Performance:\nL4 = faster, lightweight. L7 = slightly heavier due to parsing \u0026amp; TLS. Flexibility:\nL4 = great for raw TCP services (databases, game servers). L7 = ideal for APIs, web apps, microservices (content-based routing, TLS offload, authentication, caching). Design choice boils down to speed vs flexibility.\nPrototype: HAProxy + Python Flask Server # Prototype Link You can build a simple prototype to experiment with:\nHAProxy as the L7 load balancer. Flask running multiple backend servers (/api, /images). HAProxy routes traffic based on path rules. This setup helps visualize how requests flow through a real-world load balancing pipeline.\nFinal Thoughts # Load balancers may be invisible, but they are critical for modern distributed systems.\nChoosing between L4 vs L7 comes down to whether you need performance (L4) or intelligence (L7).\nThe best architectures often combine both, with L4 for raw speed and L7 for smart routing \u0026amp; app awareness.\n","date":"17 August 2025","externalUrl":null,"permalink":"/blogs/load_balancers/","section":"Blogs","summary":"","title":"Load Balancers: The Invisible Backbone of Modern Applications","type":"blogs"},{"content":" Graduate student passionate about operating systems, computer networking, and distributed systems with a focus on building high-performance, scalable, and fault-tolerant systems. Hands-on experience in developing features for CRM platforms, contributing to seamless user experiences and platform optimization. # San Jose State University 2024 - current MS in Computer Software Engineering Developed an autonomous AI support agent that answers customer queries from help docs and monitors user emotions to flag poor experiences. Escalates to human agents when negative sentiment is detected. Designed a private, distributed blockchain system to ensure the integrity of file access audits using a PoA-inspired consensus mechanism. Developed a scalable social platform for networking and job search, utilizing a queue to handle notification burst scaling and a cache for optimizing database query performance. Built a high-performance distributed document query engine with focus on scalability, fault tolerance, and low-latency response. Zoho Corporation 2018 - 2024 Member Technical Staff Led the major CRM+ release by managing safe data migrations and feature rollouts to ensure a stable, user-friendly upgrade experience. Directed end-to-end development of CRM features—Budget Board, Marketing Calendar, and GDPR tools—leading a 5-member team to serve 250K+ businesses. Improved mobile experience by enabling deep linking from web to app views and standardizing APIs for smoother service integration. Coordinated security efforts by enforcing security guidelines, triaging bug reports, and sharing RCAs to improve team-wide secure coding practices. ","date":"9 August 2025","externalUrl":null,"permalink":"/about/","section":"","summary":"","title":"About Me","type":"page"},{"content":" View GitHub Profile AI-Driven Kubernetes Observability Platform # Built an autonomous diagnostic system using LangGraph, Prometheus, Loki, Tempo, and Grafana MCP for cross-signal querying, enabling automated root-cause analysis and reducing MTTR Engineered a multi-agent diagnosis pipeline with integrated tracing (Langfuse) and automated regression evaluation workflows (Weights and Biases), establishing reproducible performance benchmarks and improving agent-level observability. Source Code AI Agents AWS EKS Amazon SQS Langraph Claude Code Into the Wild # Built a cross-platform AI survivalist chat application using Flutter and FastAPI, supporting web and Android users with real-time group chat, Google SSO, speech-to-text, image input, and AI-powered plant, animal, and survival-question assistance. Designed a cloud-native RAG and MLOps pipeline using iNaturalist and Wikipedia data, AWS Bedrock/Cohere embeddings, Pinecone, Gemini, Docker, Kubernetes on AWS EKS, and Terraform, with security, observability, and LLM evaluation through Cloudflare, Firebase, and Weights \u0026amp; Biases. Source Code Cohere Embeddings RAG Architecture AWS IAM Weights \u0026amp; Biases Pinecone WebSockets Google OAuth Flutter, Dart CareerCon - Social Network # Built a full-stack job search and hiring platform for applicants, companies, and moderators, supporting applicant/company profile management, job posting and search, resume/CV management, application tracking, real-time alerts, AI chatbot assistance, and employer dashboards. Designed a scalable microservices and event-driven architecture using React, REST APIs, Auth0 SSO/2FA, MongoDB, Redis, RabbitMQ, Nginx, GitHub Actions, Tableau, and an LLM-powered RAG pipeline with LangChain, ChromaDB, and Gemini for semantic job search, skill extraction, and contextual career guidance. Source Code Typescript Event Driven Architecture Microservices RabbitMQ Redis MongoDB Nginx GitHub Actions Tableau Auth0 SSO/2FA Distributed Document Query Engine # Architected a high-performance C++ and Go distributed query engine using gRPC for inter-node communication and OpenMP for parallel execution, enabling low-latency search across multi-gigabyte datasets. Optimized multi-threaded execution and memory utilization, achieving 29% faster query execution and improving CPU utilization efficiency by 35% over baseline implementations. Source Code c++ OpenMP Distributed Systems MemoryMapped IO AuditChain - Blockchain-Based Audit Logging System # Built a distributed blockchain-based audit logging system that records file access operations such as read, write, update, and delete with user metadata, timestamps, RSA digital signatures, and cryptographic verification to ensure tamper-proof audit trails.\nImplemented a private validator-node network using gRPC, Protocol Buffers, Merkle trees, SHA-256 block hashing, leader-based consensus, heartbeat monitoring, leader election, mempool processing, and two-phase block commit to provide reliable, fault-tolerant audit log validation.\nSource Code Python Asyncio Protobuf PyBind Consensus Algorithms RSA Digital Signatures Cryptographic Verification Leader Election Emotion-Aware Agentic Customer Support System # Developed an autonomous AI support agent that answers customer queries from help docs and monitors user emotions to flag poor experiences. Escalates to human agents when negative sentiment is detected.\nSource Code Agents RAG Architecture ","date":"9 August 2025","externalUrl":null,"permalink":"/projects/","section":"","summary":"","title":"Projects","type":"page"},{"content":"When it comes to speeding up data retrieval, indexes are the backbone of modern databases. Among the many indexing techniques out there, B-trees are one of the most widely used—powering systems like MySQL, PostgreSQL, and Oracle. In this post, we’ll explore what B-tree indexing is, how it works, how it compares with LSM trees, and what optimizations make it even better.\nWhat is B-Tree Indexing? # At a high level, B-tree indexing organizes key-value pairs in sorted order, which allows for fast lookups and range queries. But beyond this similarity with structures like SSTables, B-trees follow a fundamentally different design philosophy.\nPage-Based Structure # B-trees break the database into fixed-size blocks or pages, typically 4KB in size. This design closely mirrors hardware-level storage, where disk I/O also happens in fixed-size blocks. Each page can reference other pages via disk addresses—similar to pointers in memory.\nThere’s a single root page at the top of the B-tree. To find a key, the database engine starts here and traverses down the tree by following references. Each intermediate page acts like a decision tree, mapping key ranges to child pages. Eventually, the search lands on a leaf page, which either stores the values directly or references where they can be found.\nThis multi-level structure makes B-trees logarithmically scalable. With a branching factor of 500 and 4KB pages, just four levels can index up to 256 TB of data!\nModifying B-Trees: Updates and Inserts # Updating a value in a B-tree is straightforward:\nFind the leaf page containing the key Modify the value Write the page back to disk Since page locations remain stable, references to them do not change. Inserting a new key is similar—if there’s space in the correct page, it’s inserted directly. If not, the page is split in two, and the parent page is updated accordingly. This keeps the tree balanced, maintaining its O(log n) depth.\nB-Trees vs. LSM Trees # B-trees and Log-Structured Merge (LSM) Trees take opposite approaches to storage:\nFeature B-Trees LSM Trees Write behavior In-place updates Append-only Crash safety Requires WAL Built-in via immutable segments Read performance Fast for random lookups Optimized for writes, reads can be costly without compaction Concurrency Requires latches Easier to scale writes The key difference lies in their write patterns. While B-trees modify data in place, LSM trees append data to new segments. This makes B-trees more sensitive to crashes—partially written pages can lead to corruption. To prevent this, B-trees use a Write-Ahead Log (WAL), ensuring changes are safely logged before they\u0026rsquo;re applied.\nConcurrency is another area where B-trees require careful handling. Since multiple threads may access or modify the tree, B-tree implementations use latches (lightweight locks) to maintain consistency.\nOptimizations: B+ Trees and Beyond # Modern systems often use B+ Trees, a variant of B-trees with additional enhancements:\nLeaf pages include pointers to siblings, enabling fast in-order scans without backtracking. All values are stored in the leaf nodes, keeping internal nodes smaller and more efficient. Some systems even hybridize B-trees and log-structured concepts. For example, Fractal Trees incorporate buffering techniques from LSM trees to reduce disk seeks while preserving the page-oriented structure of B-trees.\nFinal Thoughts # B-tree indexing strikes a balance between fast reads and manageable writes, making it ideal for a wide range of workloads. While log-structured approaches like LSM trees have grown in popularity for write-heavy applications, B-trees remain the go-to choice for OLTP systems that demand low-latency lookups and fine-grained updates.\nUnderstanding B-trees helps you make informed decisions when choosing a database engine or tuning its performance—and that’s the kind of knowledge that pays off at scale.\nReferences # B-Tree Visualize ","date":"9 August 2025","externalUrl":null,"permalink":"/blogs/btree_index/","section":"Blogs","summary":"","title":"Demystifying B-Tree Indexing in Databases","type":"blogs"},{"content":"Modern storage engines like RocksDB, LevelDB, and Apache Cassandra rely on powerful data structures to handle large volumes of data efficiently. Two such core concepts are SSTables (Sorted String Tables) and LSM Trees (Log-Structured Merge Trees). In this post, we’ll walk through what they are, how they work together, and how they support scalable storage systems.\nWhat Are SSTables? # An SSTable (Sorted String Table) is a file format that stores a sorted sequence of key-value pairs. The key idea is that all keys are ordered, and each key appears only once per file.\nWhy Use SSTables? # Compared to unordered log segments (e.g., with hash indexes), SSTables offer several key benefits:\nEfficient Merging: Sorted data allows merging multiple files using a simple and fast merge sort algorithm.\nSparse Index for Faster Lookups: Because keys are sorted, you don’t need a full in-memory index—only a sparse index storing selected key offsets. This reduces memory usage while allowing fast lookups.\nBlock Compression: Group adjacent key-value pairs into compressed blocks. The sparse index can point to the start of each compressed block, improving I/O efficiency and reducing disk space usage.\nReal-World Usage # Apache Cassandra stores its immutable SSTables on disk and relies heavily on their sorted nature for compaction and querying.\nRocksDB and LevelDB use SSTables to persist sorted key-value pairs during memtable flushes.\nWhat Are LSM Trees? # A Log-Structured Merge-Tree (LSM Tree) is a write-optimized data structure that organizes multiple SSTables in levels and merges them gradually in the background.\nKey Properties # High Write Throughput: Writes are first buffered in memory and then flushed sequentially to disk, making disk I/O highly efficient.\nSorted Storage: Enables efficient range queries without needing to load the entire dataset into memory.\nBackground Compaction: Merges SSTables, removes old/deleted entries, and keeps the number of SSTables manageable.\nReal-World Usage # LevelDB (by Google): A lightweight embedded key-value store built on LSM Trees.\nRocksDB (by Meta): A high-performance fork of LevelDB optimized for SSDs, also based on LSM Trees.\nApache HBase: Uses LSM Trees to manage large-scale time-series and columnar data.\nApache Cassandra: Combines LSM Trees with distributed replication and tunable consistency.\nHow SSTables and LSM Trees Work Together? # Here’s how a typical storage engine uses SSTables in an LSM Tree architecture:\nMemtable Writes: Incoming writes go into an in-memory structure called a memtable.\nFlush to SSTables: When the memtable exceeds a size threshold, it’s flushed to disk as a new SSTable.\nHierarchical Reads: Read requests check:\nMemtable\nThe newest SSTables\nOlder SSTables, in sequence\nCompaction: SSTables are periodically merged and compacted in the background to optimize space and remove outdated data.\nOptimizations in LSM Trees # Write-Ahead Logging (WAL) To prevent data loss on crashes, every write is first recorded in a write-ahead log before being added to the memtable.\nBloom Filters To avoid expensive lookups for non-existent keys, each SSTable maintains a Bloom filter, which tells us with high probability if a key is not in the file—saving disk I/O.\nUsed by: RocksDB, LevelDB, Cassandra\nCompaction Strategies Two widely used strategies: Size-Tiered Compaction (STC): Merge SSTables of similar size together.\nLeveled Compaction (LCS): Organize SSTables into levels where each level holds files of exponentially increasing size and tighter key range bounds.\nRocksDB uses leveled compaction by default for better read performance.\nFinal Thoughts # SSTables and LSM Trees are at the heart of modern high-performance databases. Their elegant combination of write efficiency, sorted storage, and compaction make them perfect for workloads that demand high throughput and scalable reads.\nUnderstanding these concepts is essential whether you’re tuning a database like RocksDB, building on top of Cassandra, or designing your own storage layer.\n","date":"8 August 2025","externalUrl":null,"permalink":"/blogs/lsmtrees/","section":"Blogs","summary":"","title":"SSTables and LSM Trees in Modern Storage Engines","type":"blogs"},{"content":" Hash Indexing: A Foundation for Key-Value Storage # Hash indexing is a method used to index key-value data. It is one of the simplest and most fundamental indexing strategies, forming the basis for more complex structures like B-trees and LSM-trees.\nKey-Value Indexing and Hash Functions # Key-value stores are quite similar to the dictionary or HashMap data types found in most programming languages, typically implemented using hash functions.\nSimple Idea # Indexing: Maintain an in-memory hash map that maps each key to a byte offset in the data file—indicating where the value is stored. File Storage: Use append-only files (as introduced in the previous blog post). Example:\nIf the file contains:\n0 url1 data1 25 url2 data2 50 url3 data3 Then the hash map looks like:\n\u0026#34;url1\u0026#34; → 0 \u0026#34;url2\u0026#34; → 25 \u0026#34;url3\u0026#34; → 50 How to Scale Append-Only Storage? # Introducing: Data Compaction # To avoid endlessly growing files, we segment the logs and periodically merge them to remove duplicates and deleted (tombstoned) entries.\nWhat Is Data Compaction? # Store logs in segments The last segment is the most recent Periodically merge segments: Discard outdated or deleted keys Retain only the latest values Example:\nSegment 1: \u0026#34;url1\u0026#34;: \u0026lt;data1\u0026gt; \u0026#34;url2\u0026#34;: \u0026lt;data2\u0026gt; \u0026#34;url3\u0026#34;: \u0026lt;data3\u0026gt; Segment 2: \u0026#34;url1\u0026#34;: \u0026lt;TOMBSTONE\u0026gt; \u0026#34;url2\u0026#34;: \u0026lt;new_data2\u0026gt; \u0026#34;url4\u0026#34;: \u0026lt;data4\u0026gt; After compaction: \u0026#34;url2\u0026#34;: \u0026lt;new_data2\u0026gt; \u0026#34;url3\u0026#34;: \u0026lt;data3\u0026gt; \u0026#34;url4\u0026#34;: \u0026lt;data4\u0026gt; The compaction process runs in a background thread, allowing the main thread to continue serving requests without blocking.\nIn-Memory Segment Hash Maps # Each segment maintains its own in-memory hash map.\nWhen looking up a key:\nFirst check the most recent segment’s hash map If not found, check older segments in order This design keeps recent data retrieval fast while gradually cleaning up old data in the background.\nWhy Use Append-Only File Systems? # Faster Writes: Appending and merging are sequential operations, faster than random writes (especially on HDDs, and still favorable on SSDs). Simpler Crash Recovery: With immutable segments, you avoid partial overwrites or corrupted files. Reduces Fragmentation: Periodic compaction prevents file fragmentation and stale data buildup. What Are the Trade-Offs? # No Efficient Range Queries: Hash indexes do not preserve key order You can\u0026rsquo;t easily scan from kitty00000 to kitty99999 Memory Limitations: Hash table must fit entirely in memory for good performance Disk-based hash maps are difficult to scale On-Disk Hash Maps Are Problematic: Random I/O makes lookups slow Resizing is expensive Collision handling is complex Example # Suppose you have 10 billion keys—too large for RAM—so you store the hash table on disk.\nLookup: To find \u0026#34;alice@example.com\u0026#34;, you must calculate its hash and jump to the right disk offset. This causes multiple random I/O operations Updates and inserts are also inefficient Hash tables are designed for fast, random access—but only when the full structure fits in memory. On disk, they become slow and hard to maintain.\nFinal Thoughts # Hash indexes are a powerful choice for exact key lookups in memory-efficient scenarios. However, their limitations with range queries and disk access pave the way for more advanced structures like SSTables, LSM-trees, and B-trees.\nIn the next post, we’ll explore those alternatives—and how they build on the simplicity of hash indexing to scale better with large datasets.\n","date":"7 August 2025","externalUrl":null,"permalink":"/blogs/hash_index/","section":"Blogs","summary":"","title":"Hash Indexing","type":"blogs"},{"content":" Why Do We Need Indexing? # Let\u0026rsquo;s Implement the Simplest Key-Value Store Using Bash and a File # These two functions implement a very basic key-value store using Bash:\n# Append a key-value pair to the database file db_set(){ echo \u0026#34;$1 $2\u0026#34; \u0026gt;\u0026gt; database } # Get the most recent value for a given key db_get(){ grep \u0026#34;^$1 \u0026#34; database | tail -n 1 | awk \u0026#39;{print $2}\u0026#39; } This is an example of a log-structured or append-only data file. Each time we write a new value, we simply append to the end.\nThe Problem with Lookups # Every time you want to look up a key using db_get, it has to scan the entire file from beginning to end, searching for all occurrences of the key. It then picks the most recent one.\nIn algorithmic terms, the cost of a lookup is O(n) — that is, the time taken grows linearly with the number of records.\nFor example, if the file has 1,000 entries, it may need to check all 1,000 lines to find the correct value.\nWhat’s the Solution? Indexing! # To efficiently find the value for a particular key, we need a different structure: an index.\nAn index is additional metadata that acts as a signpost, helping you quickly locate the data you want — without scanning the entire dataset.\nHow Does an Index Work? # Let’s say our data file looks like this:\ndatabase: A 10 B 20 C 30 An index might look like:\nA → line 1 B → line 2 C → line 3 So instead of scanning line-by-line, we can jump directly to the location for key B.\nIn real-world databases, indexes are much more sophisticated, but the principle remains the same.\nWhat’s the Trade-Off? # Indexes greatly improve read performance, but they come with costs:\nSlower writes — because the index also needs to be updated every time you write data. Extra storage — because indexes consume space. That’s why most databases don’t index everything by default. Instead, they let you — the developer or DBA — decide which columns to index based on query patterns.\nWhat’s Next? # In the next part of this series, we’ll explore:\nHow different types of indexes (e.g., Hash, B-Tree) work How real databases maintain indexes efficiently during updates and deletions Stay tuned!\n","date":"6 August 2025","externalUrl":null,"permalink":"/blogs/indexing/","section":"Blogs","summary":"","title":"Intro to Indexing","type":"blogs"},{"content":"","externalUrl":null,"permalink":"/authors/","section":"Authors","summary":"","title":"Authors","type":"authors"},{"content":"","externalUrl":null,"permalink":"/categories/","section":"Categories","summary":"","title":"Categories","type":"categories"},{"content":"","externalUrl":null,"permalink":"/posts/","section":"Posts","summary":"","title":"Posts","type":"posts"},{"content":"","externalUrl":null,"permalink":"/series/","section":"Series","summary":"","title":"Series","type":"series"},{"content":"","externalUrl":null,"permalink":"/tags/","section":"Tags","summary":"","title":"Tags","type":"tags"},{"content":"AI agents are becoming a big deal in the world of technology. What makes them special? Unlike traditional softwares, These are programs that aren’t stuck following one fixed path—they’re built to be flexible. Instead of hard coding what steps to take, they figure things out as they go, depending on the context.\nHow is this possible? # This flexibility comes from big improvements in how AI models predict the “next word” (or token) in a sentence, based on what’s already been said. That might sound simple, but it’s opened up a whole new world where these models can make surprisingly smart decisions.\nAt the heart of these AI agents is something called a large language model (LLM). You can think of it as the agent’s brain. It helps the agent understand what’s going on, figure out what tools it has, and decide what steps to take—without needing someone to guide it every time. In this blog, we’ll explore how these agents work, what makes them special, and why they’re changing the game.\nHow AI Agents Use Language Models ? # At the core of every smart AI agent is something called a large language model (LLM)—basically, a very powerful autocomplete engine. LLMs don’t think like humans, but they’re really good at guessing the next word (or piece of a word), one tiny step at a time. These tiny pieces are called tokens, and predicting them well is what helps an agent seem smart or helpful in a conversation.\nTo keep track of what’s happening in a chat or task, LLMs use something called attention. This lets them “remember” important parts of what’s been said, so they can stay on topic and respond in a useful way. It’s like paying attention to the right parts of a conversation to make a good reply.\nWe talk to LLMs using something called prompting, which means giving them input in a special format they understand. These prompts include all the roles and messages in the conversation—like who’s the user, who’s the assistant, and what’s already been said. For example:\nconversation = [ {\u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;I need help with my order\u0026#34;}, {\u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;I\u0026#39;d be happy to help. Could you provide your order number?\u0026#34;}, {\u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;It\u0026#39;s ORDER-123\u0026#34;}, ] Under the hood, it might look more like this:\n\u0026lt;|im_start|\u0026gt;system You are a helpful AI assistant named SmolLM, trained by Hugging Face\u0026lt;|im_end|\u0026gt; \u0026lt;|im_start|\u0026gt;user I need help with my order\u0026lt;|im_end|\u0026gt; These special tokens (like \u0026lt;|im_start|\u0026gt;) help the model understand who’s speaking and what its role is. By feeding in a prompt like this, the model can decide how to respond—and that’s how agents “think” and talk.\nWhat Are Tools (and Why Do Agents Need Them)? # While LLMs are great at understanding and generating text, they’re not perfect at everything. Sometimes, they need help—especially when it comes to doing things outside their core skill set, like solving math problems, checking live data, or interacting with other systems. That’s where tools come in.\nA tool is basically an extra ability that you give the AI. Think of it like handing a calculator to someone who’s good with words but not great with numbers. The model still drives the conversation, but when it realizes it needs to calculate something, it can call on the calculator tool to do it right.\nAnother big reason tools are important is that LLMs only know what they were trained on. That means they don’t know what happened in the world after their training finished. So, if your agent needs the current weather, stock prices, or news headlines, you’ll need to give it a tool that can look those things up.\nTo make this work, we also need to explain each tool to the LLM ahead of time—kind of like giving it instructions. For example, you might say: “Here’s a weather tool. Use it when someone asks about the current weather.” Once the model knows what the tool is for, it will decide on its own whether to use it based on the situation.\nSo in short: tools make agents smarter, more useful, and more connected to the real world.\nHow Do Agents Think and Make Decisions? # To act on their own, AI agents need a way to break down problems, take steps, and figure out when they’ve reached a goal. One popular prompting technique for this is called the Thought-Action-Observation (TAO) cycle. It guides the agent to follow a loop: think about the situation, take an action (like calling a tool), observe what happened, and then do it again—until the task is complete.\nIf you’ve done any programming, it’s kind of like a while loop while (task not complete) → think → act → observe → repeat\nHere’s how a prompt might change during each step of the cycle:\nUser: What’s the weather in New York and do I need an umbrella? \u0026lt;THINK\u0026gt; I need to get the current weather in New York to answer this. \u0026lt;/THINK\u0026gt; \u0026lt;ACTION\u0026gt; call_weather_api(location=\u0026#34;New York\u0026#34;) \u0026lt;/ACTION\u0026gt; \u0026lt;OBSERVATION\u0026gt; The forecast says rain with 80% chance. \u0026lt;/OBSERVATION\u0026gt; \u0026lt;THINK\u0026gt; Since it\u0026#39;s going to rain, the user will likely need an umbrella. \u0026lt;/THINK\u0026gt; \u0026lt;FINAL_ANSWER\u0026gt; Yes, you should take an umbrella. \u0026lt;/FINAL_ANSWER\u0026gt; This cycle helps the agent slow down, reason through the task, and use tools as needed—without jumping straight to a guess.\nMore recently, some advanced models like DeepSeek-V2 or OpenAI’s o1 are taking this even further. They’ve been fine-tuned to think before answering, using structured tags like and to explicitly separate the reasoning phase from the final response. Unlike TAO, which is just a clever way to craft the prompt, this is a training-level technique, where the model actually learns to reason step by step during training.\nSo in short: TAO is a prompting method that helps any LLM think more like an agent. But newer models are starting to learn that behavior right out of the box.\nWrapping Up # Building AI agents might sound magical, but under the hood, there’s a lot going on. You need to design smart prompts, manage the agent’s state as it moves through tasks, handle tool calls, and keep everything in sync across multiple steps. It can get messy fast.\nThankfully, you don’t have to do all of this by hand. Frameworks like SmolAgents and LangGraph take care of much of the heavy lifting. They help you build structured, reliable agents without reinventing the wheel every time.\nWe’ll explore how these frameworks work—and how you can use them to build your own AI agents—in the next blog. Stay tuned!\nReferences:\nIntro to agents by Hugging face Deep Learning course by 3Blue1brown ","externalUrl":null,"permalink":"/blogs/agents_intro/","section":"Blogs","summary":"","title":"What are AI Agents ?","type":"blogs"},{"content":"Large Language Models (LLMs) are impressively smart—they can write essays, explain complex topics, even draft code. But when it comes to answering questions based on the latest information or your private data, they run into some pretty big roadblocks.\nSo, what’s the problem? # They can’t access recent data\nMost LLMs are trained on data that cuts off at a specific point (for example, September 2021 or January 2023). If something happened after that, they have no idea.\nRetraining is hard (and expensive)\nUpdating a model with new information means retraining it, which is resource-heavy, time-consuming, and not practical to do frequently—especially if your data changes every week or day.\nHallucinations\nSometimes LLMs \u0026ldquo;make things up\u0026rdquo;—they generate answers that sound right but aren’t based on any real data. This can be risky in areas like healthcare, law, or internal business use.\nThis is where RAG comes in # RAG (Retrieval-Augmented Generation) is a smart workaround to all these issues. Instead of trying to retrain the model every time your data changes, RAG gives the model a way to look things up on the fly.\nHere’s how it works in simple terms:\nYou store your actual data (documents, notes, PDFs, internal wikis, etc.) in a way the model can search through. When you ask a question, the system first retrieves relevant pieces of that data. Then the language model uses those pieces to generate a response—grounded in real, up-to-date information. It’s like giving the model a set of cheat sheets to refer to before answering.\nChunking: Breaking Down Data the Smart Way # Before feeding your documents into a Retrieval-Augmented Generation (RAG) system, there’s one crucial step: chunking.\nLLMs have a limit on how much text they can “see” at once—this is called the token limit. If your documents are too big, they can’t all fit in the model’s context window. That’s where chunking comes in: we split large documents into smaller pieces, or \u0026ldquo;chunks\u0026rdquo;, that the model can handle.\nBut here’s the tricky part—how you chunk the data really matters.\nWhy chunking needs balance # Large chunks can go over the token limit or overwhelm the model. Important parts in the middle might get lost or ignored. Tiny chunks are easier to fit, but can lose context and meaning, making it harder to generate accurate answers. The sweet spot? It depends on the nature of your data. A technical manual might work well split by sections or headers, while chat logs may need to be grouped by conversation turns.\nDifferent ways to chunk # There’s no one-size-fits-all approach. Some common methods include:\nCharacter-based splitting: Break by a set number of characters. Sentence-based splitting: Break on sentence boundaries for more natural chunks. Semantic splitting: Use meaning and structure (via embeddings or ML models) to split at logical points in the text. If you\u0026rsquo;re curious to explore different chunking strategies in practice, here\u0026rsquo;s a fantastic tutorial showing five levels of text splitting:\n5 Levels of Text Splitting - RetrievalTutorials on GitHub Chunking may seem like a small step, but it plays a huge role in the accuracy and efficiency of your RAG pipeline.\nWhat Are Embeddings? # Once your data is chunked, the next step is to make it searchable for your RAG system. This is where embeddings come in.\nThink of embeddings as a way to turn words, sentences, or entire documents into numbers—specifically, vectors (which are just long lists of numbers). These numbers capture the meaning of the text, so that similar chunks end up with similar embeddings.\nWhy embeddings matter # You can’t compare two pieces of text directly in raw form (e.g., “How to cook rice” vs. “Steps to boil white rice”), but once they’re turned into embeddings, the system can measure how similar they are—kind of like measuring distance between points on a map.\nSo when you ask a question like “How do I cook rice?”, the RAG system can search through your embedded chunks and find the ones closest in meaning—even if the wording is totally different.\nModels Used to Create Embeddings # There are several popular models that generate embeddings, and each comes with trade-offs in terms of quality, speed, and cost.\nHere are a few common ones:\n🔹 OpenAI’s text-embedding-3-small or text-embedding-ada-002 # Pros: High-quality, fast, easy to integrate. Cons: Paid API, so cost can add up for large-scale use. 🔹 HuggingFace models (e.g. all-MiniLM-L6-v2) # Pros: Open-source, free to run locally or on your own servers. Cons: Slightly lower performance than OpenAI for some tasks; requires some setup. 🔹 Cohere, Google, and others # Pros: Strong performance; often tuned for different use cases (e.g., longer texts). Cons: Varies by provider—some are paid, some have limitations. Indexing Embeddings and FAISS # So now we’ve got our text chunks turned into embeddings—cool! But what happens when a user asks a question? How does the system quickly find the most relevant chunks from possibly thousands (or millions) of embeddings?\nThat’s where indexing comes in.\nWhy Indexing Matters # Imagine you’ve got a massive list of coordinates (your embeddings), and someone asks: “Which of these is closest to this new point?”\nYou could check every single one, but that would be slow and inefficient—especially as your dataset grows.\nInstead, we use special data structures to speed this up. These structures are called indexes, and they’re built to make similarity search fast and scalable.\nMeasuring Similarity: Distance Matters # To find similar embeddings, we calculate the distance between them. The closer two vectors are, the more similar the texts they represent.\nThere are a few common distance types:\nCosine Similarity: Measures the angle between two vectors (ignores length). Often used when vectors are normalized. Euclidean Distance: Measures the straight-line distance between two points. Dot Product: Used in some models to capture both similarity and magnitude. Normalization (scaling vectors to unit length) helps make comparisons more stable—especially when using cosine similarity. It ensures you\u0026rsquo;re comparing direction, not size.\nTypes of Indexes # Depending on your needs (speed vs. accuracy), you can choose different types of indexes:\nFlat Index: Brute-force, checks every vector. Very accurate, but slow for large datasets. IVF (Inverted File Index): Groups vectors into clusters—faster but slightly less accurate. HNSW (Hierarchical Navigable Small World): Graph-based structure, great for high speed and high accuracy. PQ (Product Quantization): Compresses vectors to save space—useful for massive datasets. Enter FAISS: Facebook AI Similarity Search # FAISS is an open-source library by Facebook AI that makes all of this practical.\nBuilt for fast vector search, even with millions of embeddings. Supports multiple index types (Flat, IVF, HNSW, PQ, and combos). Can run on CPU or GPU, and works great with Python. Easy to integrate into your RAG pipeline. With FAISS, your system can instantly retrieve the most relevant text chunks, even from a massive collection, making your RAG responses snappy and accurate.\nVector Databases: Managing Embeddings at Scale # Once you\u0026rsquo;ve got embeddings and an index (like FAISS), you need a way to store, manage, and search them effectively—especially in real-world applications where your data changes, grows, or needs to be queried alongside metadata.\nThis is where vector databases come in.\nWait—Can’t We Just Use a Regular Database? # Traditional databases (like PostgreSQL, MongoDB, MySQL) are amazing for structured data—think rows, columns, and filters like age \u0026gt; 30. But they weren’t designed for vector math or similarity search.\nHowever, things are changing. Many traditional databases are now adding vector support:\nPostgreSQL has extensions like pgvector, which allow you to store and search vector embeddings. MongoDB introduced native vector search in recent versions. Pros:\nFamiliar ecosystem and tools Combine structured filters (e.g., tags, categories) with vector search Easier for teams already using these databases Cons:\nMay not scale well for very large vector datasets Slower or less optimized compared to dedicated vector DBs Dedicated Vector Databases # These are built from the ground up to handle vector data—fast search, scaling, and features like hybrid search (vector + keyword). Here are some of the most popular ones:\n🔹 Pinecone # Pros: Fully managed, fast, supports metadata filtering, scales well Cons: Paid service; less control if you need full self-hosting 🔹 Weaviate # Pros: Open-source or managed, supports semantic and hybrid search, auto chunking, metadata-aware Cons: Heavier footprint for self-hosting; steeper learning curve 🔹 Milvus # Pros: High performance, GPU acceleration, large-scale production-ready Cons: Needs more infra to set up; more suitable for enterprise use 🔹 Qdrant # Pros: Lightweight, easy API, fast, open-source Cons: Fewer integrations than some competitors (but improving fast) 🔹 Vespa, Vald, Zilliz (Milvus-backed) — also great choices depending on your use case and infra preferences # So… Which One Should You Use? # Just getting started? Try pgvector with PostgreSQL or Qdrant for a simple, open-source experience. Scaling up? Look at Pinecone, Weaviate, or Milvus. Need full control? Go for open-source + FAISS or a self-hosted vector DB. In short, vector databases are the backbone of RAG at scale. They make it possible to store thousands or millions of embeddings and search them in milliseconds—often combined with filters like document type, source, or time.\nConclusion: RAG Is Just the Beginning # RAG—Retrieval-Augmented Generation—is a powerful way to make LLMs smarter, more reliable, and grounded in real data. By combining a model’s ability to generate text with the ability to look things up, we overcome a lot of the core limitations of standalone LLMs.\nWe’ve walked through the core steps:\nChunking your data the right way to balance context and size Embedding those chunks into a form machines can understand Indexing them for fast and efficient search Using vector databases to store and manage it all at scale But that’s not the whole story.\nThere are plenty of advanced techniques that can push RAG even further:\nRe-ranking: After the initial retrieval, you can use another model to sort or refine the top results for better relevance. Task Decomposition: You can use one LLM to break down a complex question into smaller tasks, and have another LLM (or the same one) answer each with more precision. Agents \u0026amp; Recursive Retrieval: Intelligent agents can dig deeper—asking follow-up questions, retrieving more info in multiple rounds, and stitching it all together to give better answers. These strategies make RAG systems more dynamic, more accurate, and more adaptable—especially in complex domains like legal, medical, or enterprise use cases.\nReferences : # Advance RAG - Huggingface cookbook Augmented Language Models ","externalUrl":null,"permalink":"/blogs/rag/","section":"Blogs","summary":"","title":"What is RAG ?","type":"blogs"}]