<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Praveen Labs]]></title><description><![CDATA[Praveen Labs]]></description><link>https://dev-praveen.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Praveen Labs</title><link>https://dev-praveen.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Tue, 06 Oct 2026 13:06:58 GMT</lastBuildDate><atom:link href="https://dev-praveen.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Exploring GenAI Retrieval with a Resume: Keyword, Semantic, and Hybrid Search]]></title><description><![CDATA[I’ve been spending some time learning and exploring Generative AI concepts, especially around embeddings, retrieval, and RAG.
Instead of starting with a large production-style application, I wanted to]]></description><link>https://dev-praveen.hashnode.dev/exploring-genai-retrieval-with-a-resume-keyword-semantic-and-hybrid-search</link><guid isPermaLink="true">https://dev-praveen.hashnode.dev/exploring-genai-retrieval-with-a-resume-keyword-semantic-and-hybrid-search</guid><category><![CDATA[generative ai]]></category><category><![CDATA[genai]]></category><category><![CDATA[RAG ]]></category><category><![CDATA[hybrid search]]></category><category><![CDATA[semantic search]]></category><category><![CDATA[keyword search]]></category><category><![CDATA[#Embeddings]]></category><category><![CDATA[Document Search]]></category><dc:creator><![CDATA[praveenkulharee]]></dc:creator><pubDate>Sun, 04 Oct 2026 12:40:07 GMT</pubDate><content:encoded><![CDATA[<p>I’ve been spending some time learning and exploring Generative AI concepts, especially around <strong>embeddings, retrieval, and RAG</strong>.</p>
<p>Instead of starting with a large production-style application, I wanted to understand one of the fundamental parts of a RAG system:</p>
<blockquote>
<p><strong>How does a system decide which piece of a document is relevant to a user's question?</strong></p>
</blockquote>
<p>So I built a small personal learning project around something familiar: <strong>my resume</strong>.</p>
<p>The goal wasn't to build a production resume parser or commercial AI application. It was simply to learn by building and experimenting.</p>
<p>The project takes a resume from a <code>.docx</code> file, extracts the content, splits it into chunks, generates embeddings, and allows me to experiment with different retrieval approaches:</p>
<ul>
<li><p>Keyword search</p>
</li>
<li><p>Semantic search</p>
</li>
<li><p>Hybrid search</p>
</li>
</ul>
<hr />
<h1>The basic pipeline</h1>
<p>The overall flow looks like this:</p>
<p><strong>Resume (.docx)</strong><br />↓<br /><strong>Text extraction</strong><br />↓<br /><strong>Chunking</strong><br />↓<br /><strong>Embeddings</strong><br />↓<br /><strong>Keyword / Semantic / Hybrid search</strong><br />↓<br /><strong>Relevant chunks</strong></p>
<p>The resume content is extracted from paragraphs and tables and split into overlapping chunks. In this experiment, I used roughly <strong>500 words per chunk with around 100 words of overlap</strong>.</p>
<p>For semantic retrieval, I used:</p>
<p><code>BAAI/bge-small-en-v1.5</code></p>
<p>The embedding size is <strong>384 dimensions</strong>.</p>
<p>For this learning project, the vectors are kept in memory rather than using a production vector database. The focus was understanding retrieval rather than building infrastructure.</p>
<hr />
<h1>Three ways to search</h1>
<p>This was the most interesting part of the experiment.</p>
<p>Keyword, semantic, and hybrid search use different signals when deciding which chunk is relevant.</p>
<h2>1. Keyword Search — Exact terms matter</h2>
<p>Keyword search focuses on the actual words or tokens present in the query and document.</p>
<p>For example, if a resume contains:</p>
<blockquote>
<p>Python, JavaScript, Azure, Docker</p>
</blockquote>
<p>and the user searches for:</p>
<blockquote>
<p>Python</p>
</blockquote>
<p>an exact occurrence of <code>Python</code> is strong evidence that the chunk is relevant.</p>
<p>For a resume, keyword search can be useful for:</p>
<ul>
<li><p>Skills</p>
</li>
<li><p>Technology names</p>
</li>
<li><p>Company names</p>
</li>
<li><p>Email addresses</p>
</li>
<li><p>IDs</p>
</li>
<li><p>Section titles</p>
</li>
</ul>
<p>The simple keyword implementation in my project tokenizes the query and checks how many query tokens occur in each chunk.</p>
<p><strong>Query terms → matching terms → ranked chunks</strong></p>
<p>The trade-off is that keyword search doesn't naturally understand that different words can have the same meaning.</p>
<p>For example:</p>
<blockquote>
<p>"current company"</p>
</blockquote>
<p>and</p>
<blockquote>
<p>"current organization"</p>
</blockquote>
<p>may be conceptually related, but exact keyword matching doesn't automatically understand that relationship.</p>
<h2>2. Semantic Search— Meaning matters</h2>
<p>Semantic search takes a different approach.</p>
<p>Instead of asking only:</p>
<blockquote>
<p>"Does this exact word appear?"</p>
</blockquote>
<p>it asks:</p>
<blockquote>
<p>"Which piece of text has a meaning most similar to this query?"</p>
</blockquote>
<p>This is where <strong>embeddings</strong> come into the picture.</p>
<p>The document chunks are converted into vectors, and the user's query is also converted into a vector.</p>
<p>The system then compares the query vector with the chunk vectors, commonly using <strong>cosine similarity</strong>.</p>
<pre><code class="language-plaintext">Text → Embedding → Vector

Query vector ↔ Chunk vectors
</code></pre>
<p>This means the wording doesn't always have to be identical.</p>
<p>For example:</p>
<blockquote>
<p>"What technologies do you work with?"</p>
</blockquote>
<p>could be semantically related to:</p>
<blockquote>
<p>"Skills: Python, JavaScript, Azure, Docker..."</p>
</blockquote>
<p>even though the wording isn't exactly the same.</p>
<p>Semantic search is therefore useful for:</p>
<ul>
<li><p>Meaning</p>
</li>
<li><p>Context</p>
</li>
<li><p>Similar concepts</p>
</li>
<li><p>Different ways of asking the same question</p>
</li>
<li><p>Synonyms and paraphrasing</p>
</li>
</ul>
<p>The quality of this retrieval also depends on the embedding model.</p>
<h2>3. Hybrid Search — Meaning + exact terms</h2>
<p>Hybrid search combines the two signals.</p>
<pre><code class="language-plaintext">Keyword signal
       +
Semantic signal
       ↓
Combined relevance
       ↓
Ranked chunks
</code></pre>
<p>Keyword search provides precision around explicit terms, while semantic search provides contextual understanding.</p>
<p>For example:</p>
<blockquote>
<p>"Python skills used in cloud projects"</p>
</blockquote>
<p>contains both:</p>
<ul>
<li><p>a specific term: <code>Python</code></p>
</li>
<li><p>contextual intent: <code>skills used in cloud projects</code></p>
</li>
</ul>
<p>In my learning implementation, I experimented with:</p>
<pre><code class="language-plaintext">hybrid_score =
    0.7 × semantic_score +
    0.3 × keyword_score
</code></pre>
<p>These weights are not a universal rule. They are simply part of my experiment. In a real application, the appropriate balance would depend on the documents and queries being handled.</p>
<p>Document extraction and chunking can have a significant effect on what retrieval eventually sees.</p>
<h2>What I learned</h2>
<ol>
<li><p>Semantic search isn't just "search with AI"<br />It is a different retrieval mechanism based on embeddings and vector similarity.</p>
</li>
<li><p>Keyword search is still useful<br />Embeddings don't make keyword search obsolete. Exact terms such as technologies, company names, skills, and identifiers can be very useful retrieval signals.</p>
</li>
<li><p>Hybrid search isn't magic<br />Hybrid retrieval combines complementary signals, but the result still depends on the query, corpus, chunking, embedding quality, ranking method, and weighting.</p>
</li>
<li><p>Scores need context<br />A retrieval score tells me how the system ranked the available chunks. It doesn't automatically mean:</p>
<p>"This is definitely the correct answer."</p>
</li>
<li><p>Retrieval is worth inspecting separately<br />Before adding an LLM on top, I wanted to understand what was actually being retrieved.</p>
<p>If the wrong context is retrieved, putting an LLM after it doesn't automatically fix the retrieval problem.</p>
</li>
</ol>
<h2>Attachments / images</h2>
<p>I would include two visual attachments in the assignment:</p>
<h3>Search Strategy / Architecture</h3>
<p>Your existing infographic:</p>
<img src="https://cdn.hashnode.com/uploads/covers/69ee29c73d6a492cdd14e937/2e643aba-952e-44ab-a3c0-bc749b0d867c.png" alt="" style="display:block;margin:0 auto" />

<h3>Attachment 2 — Application UI / Retrieval Results</h3>
<p>Use the improved UI screenshot after the Three Ways to Search section, followed by your actual result screenshots.</p>
<p>This gives a nice progression:</p>
<ol>
<li><p>What I wanted to learn ↓</p>
</li>
<li><p>Retrieval architecture ↓</p>
</li>
<li><p>Three search strategies ↓</p>
</li>
<li><p>Application UI ↓</p>
</li>
<li><p>Same query → different results ↓</p>
</li>
<li><p>What I learned On</p>
</li>
</ol>
<img src="https://cdn.hashnode.com/uploads/covers/69ee29c73d6a492cdd14e937/f730a0ed-95a1-4b12-9333-9db0e81f89dc.png" alt="" style="display:block;margin:0 auto" />

<h2>Conclusion</h2>
<p>Choosing the right search strategy depends on what your application needs to understand.</p>
<ul>
<li><p><strong>Keyword Search</strong> works best when exact terms, names, IDs, or precise matches are important.</p>
</li>
<li><p><strong>Semantic Search</strong> focuses on meaning and intent, making it useful when users express the same idea in different words.</p>
</li>
<li><p><strong>Hybrid Search</strong> combines keyword and semantic approaches, helping balance exact matching with contextual understanding.</p>
</li>
</ul>
<p>In practice, there is no single search strategy that fits every use case. <strong>Keyword search provides precision, semantic search provides contextual relevance, and hybrid search brings both strengths together.</strong> The best approach depends on your data, query patterns, and the type of results your users expect.</p>
]]></content:encoded></item><item><title><![CDATA[How I Built a Real-Time Market Alert System using n8n (NSE + Price Tracking)]]></title><description><![CDATA[Introduction
In volatile markets, delayed decisions can lead to missed opportunities or losses.
While there are many tools that provide alerts and notifications, I wanted to build something lightweigh]]></description><link>https://dev-praveen.hashnode.dev/how-i-built-a-real-time-market-alert-system-using-n8n-nse-price-tracking</link><guid isPermaLink="true">https://dev-praveen.hashnode.dev/how-i-built-a-real-time-market-alert-system-using-n8n-nse-price-tracking</guid><category><![CDATA[n8n]]></category><category><![CDATA[automation]]></category><category><![CDATA[JavaScript]]></category><category><![CDATA[fintech]]></category><category><![CDATA[workflow]]></category><category><![CDATA[Workflow Automation]]></category><dc:creator><![CDATA[praveenkulharee]]></dc:creator><pubDate>Sun, 26 Apr 2026 16:18:07 GMT</pubDate><content:encoded><![CDATA[<h2>Introduction</h2>
<p>In volatile markets, delayed decisions can lead to missed opportunities or losses.</p>
<p>While there are many tools that provide alerts and notifications, I wanted to build something lightweight, customizable, and tailored to my needs.</p>
<p>So I built a <strong>real-time market monitoring system</strong> using n8n.</p>
<hr />
<h2>Problem Statement</h2>
<p>I needed:</p>
<ul>
<li><p>Real-time price tracking (Nifty, Crude Oil)</p>
</li>
<li><p>Daily important announcements (NSE)</p>
</li>
<li><p>Noise-free alerts</p>
</li>
<li><p>One single consolidated notification</p>
</li>
</ul>
<p>Most tools either:</p>
<ul>
<li><p>Send too many alerts ❌</p>
</li>
<li><p>Lack customization ❌</p>
</li>
<li><p>Or are paid ❌</p>
</li>
</ul>
<hr />
<h2>Solution Overview</h2>
<p>I designed a workflow that:</p>
<ul>
<li><p>Fetches market data from Yahoo Finance</p>
</li>
<li><p>Pulls announcements from National Stock Exchange of India</p>
</li>
<li><p>Applies filtering + logic</p>
</li>
<li><p>Sends alerts via Telegram</p>
</li>
</ul>
<hr />
<h2>Architecture</h2>
<p>n8n workflow architecture for real-time market alerts</p>
<img src="https://cdn.hashnode.com/uploads/covers/69ee29c73d6a492cdd14e937/7a5cf36a-c55e-4787-ad81-55683bf71021.png" alt="" style="display:block;margin:0 auto" />

<p>The workflow runs on a Cron schedule and uses parallel branches:</p>
<pre><code class="language-plaintext">Cron
 ├── NSE RSS Flow
 ├── Nifty Price Flow
 └── Crude Oil Flow
        ↓
     Merge
        ↓
   Final Alert
</code></pre>
<hr />
<h2>Key Features</h2>
<h3>1. Market Data Tracking</h3>
<ul>
<li><p>Nifty &amp; Brent Crude via Yahoo Finance API</p>
</li>
<li><p>Calculates % change vs previous close</p>
</li>
<li><p>Detects bullish 📈 / bearish 📉 signals</p>
</li>
</ul>
<hr />
<h3>2. NSE Announcements (First-Run Logic)</h3>
<ul>
<li><p>Fetches RSS feed</p>
</li>
<li><p>Runs only once per day</p>
</li>
<li><p>Filters out:</p>
<ul>
<li><p>Mutual Funds</p>
</li>
<li><p>ETFs</p>
</li>
<li><p>irrelevant entries</p>
</li>
</ul>
</li>
</ul>
<hr />
<h3>3. Smart Filtering</h3>
<ul>
<li><p>Threshold-based alerts (e.g. &gt;1%)</p>
</li>
<li><p>Avoids noise and duplicate alerts</p>
</li>
</ul>
<hr />
<h3>4. Aggregated Notification</h3>
<p>Instead of multiple alerts:</p>
<p>👉 Everything is merged into <strong>one structured message</strong></p>
<hr />
<h2>Alert Example</h2>
<pre><code class="language-plaintext">📊 Market Alert

NIFTY: +1.2% 📈
Crude Oil: -0.8% 📉

📰 NSE Updates:
• Company X result announced
• Merger news
</code></pre>
<hr />
<h2>Why I Chose n8n</h2>
<p>Instead of building a full microservice:</p>
<ul>
<li><p>Faster setup</p>
</li>
<li><p>Visual workflow</p>
</li>
<li><p>Easy integration</p>
</li>
<li><p>Flexible logic</p>
</li>
</ul>
<p>👉 Perfect for rapid prototyping and automation</p>
<hr />
<h2>Challenges Faced</h2>
<ul>
<li><p>JSON parsing issues in HTTP node</p>
</li>
<li><p>Telegram SSL issue in Docker</p>
</li>
<li><p>RSS feed returning limited data</p>
</li>
<li><p>Handling first-run logic correctly</p>
</li>
</ul>
<hr />
<h2>Learnings</h2>
<ul>
<li><p>Automation can replace repetitive monitoring</p>
</li>
<li><p>Simpler architecture &gt; over-engineering</p>
</li>
<li><p>Data filtering is more important than data fetching</p>
</li>
</ul>
<hr />
<h2>Conclusion</h2>
<p>This project started as a small experiment but turned into a very practical tool.</p>
<p>It helps:</p>
<ul>
<li><p>React faster to market changes</p>
</li>
<li><p>Reduce noise</p>
</li>
<li><p>Stay focused on important signals</p>
</li>
</ul>
<hr />
<h2>Resources</h2>
<ul>
<li><p>Yahoo Finance API</p>
</li>
<li><p>NSE RSS Feed</p>
</li>
<li><p>n8n workflow automation</p>
</li>
</ul>
]]></content:encoded></item></channel></rss>