Clelia Astra Bertelli
AI & ML interests
Recent Activity
Organizations
With 𝐢𝐧𝐠𝐞𝐬𝐭-𝐚𝐧𝐲𝐭𝐡𝐢𝐧𝐠 𝐯𝟏.𝟑.𝟎 (https://github.com/AstraBert/ingest-anything) you can now scrape content simply starting from URLs, extract the text from it, chunk it and put it into your favorite LlamaIndex-compatible database!🕸️
You can do it thanks to 𝗰𝗿𝗮𝘄𝗹𝗲𝗲 by Apify, an open-source crawling library for python and javascript that handles all the data flow from the web: ingest-anything then combines it with 𝗕𝗲𝗮𝘂𝘁𝗶𝗳𝘂𝗹𝗦𝗼𝘂𝗽, 𝗣𝗱𝗳𝗜𝘁𝗗𝗼𝘄𝗻 and 𝗣𝘆𝗠𝘂𝗣𝗱𝗳 to scrape HTML files, convert them to PDF and extract the text - hassle-free!😸
Check the attached code snippet if you're curious of knowing how to get started🎬
PS: Don't tell anybody, but this release also has another gem... It supports OpenAI models for agentic chunking, following the new releases of Chonkie🦛✨
If you don't want to miss out on the new features, leave us a little star on GitHub ➡️ https://github.com/AstraBert/ingest-anything
And join our discord community! ➡️ https://discord.gg/kDqHNjks
✅ Embeddings: now works with Sentence Transformers, Jina AI, Cohere, OpenAI, and Model2Vec
All powered via 𝗖𝗵𝗼𝗻𝗸𝗶𝗲’𝘀 𝗔𝘂𝘁𝗼𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴𝘀.
No more local-only limitations 🙌
✅ Vector DBs: now supports 𝗮𝗹𝗹 𝗟𝗹𝗮𝗺𝗮𝗜𝗻𝗱𝗲𝘅-𝗰𝗼𝗺𝗽𝗮𝘁𝗶𝗯𝗹𝗲 𝗯𝗮𝗰𝗸𝗲𝗻𝗱𝘀
Think: Qdrant, Pinecone, Weaviate, Milvus, etc.
No more bottlenecks🔓
✅ File parsing: now plugs into any 𝗟𝗹𝗮𝗺𝗮𝗜𝗻𝗱𝗲𝘅-𝗰𝗼𝗺𝗽𝗮𝘁𝗶𝗯𝗹𝗲 𝗱𝗮𝘁𝗮 𝗹𝗼𝗮𝗱𝗲𝗿
Using LlamaParse, Docling or your own setup? You’re covered.
Curious of knowing more? Try it out! 👉 https://github.com/AstraBert/ingest-anything
That's why today I'm excited to introduce 𝐫𝐞𝐚𝐝𝐞𝐫𝐬, the new feature of PdfItDown v1.4.0!🎉
With 𝘳𝘦𝘢𝘥𝘦𝘳𝘴, you can choose among three (for now👀) flavors of text extraction and conversion to PDF:
- 𝗗𝗼𝗰𝗹𝗶𝗻𝗴, which does a fantastic work with presentations, spreadsheets and word documents🦆
- 𝗟𝗹𝗮𝗺𝗮𝗣𝗮𝗿𝘀𝗲 by LlamaIndex, suitable for more complex and articulated documents, with mixture of texts, images and tables🦙
- 𝗠𝗮𝗿𝗸𝗜𝘁𝗗𝗼𝘄𝗻 by Microsoft, not the best at handling highly structured documents, by extremly flexible in terms of input file format (it can even convert XML, JSON and ZIP files!)✒️
You can use this new feature in your python scripts (check the attached code snippet!😉) and in the command line interface as well!🐍
Have fun and don't forget to star the repo on GitHub ➡️ https://github.com/AstraBert/PdfItDown
I am working on supporting compatibility with other embeddding models, and we will have that soon, for now I had to reduce the compatibility only to Sentence Transformers.
For what concerns page numbers, I am also working toward having better and more extensive metadata: everything is a big work-in-progress and will come in future releases!
So, there are two possibilities:
- If you mean customizing the embedder among the ones available within Sentence Transformers, it is very possible, you just have to change the
embedding_modelparameter when calling theingestmethod - If you mean that you have your own embedding model (like saved on your PC), that is a tad more difficult. I think Sentence Transformer might allow loading the model from your PC as long as it is compatible with the package. I think that this guide might be useful in that regard
For now the package only supports Sentence Transformers models, in the future it will probably extend its support to other embedding models as well :)
What if I told you that you can do it within three to six lines of code?🤯
Well, with my latest open-source project, 𝐢𝐧𝐠𝐞𝐬𝐭-𝐚𝐧𝐲𝐭𝐡𝐢𝐧𝐠 (https://github.com/AstraBert/ingest-anything), you can take all your non-PDF files, convert them to PDF, extract their text, chunk, embed and load them into a vector database, all in one go!🚀
How? It's pretty simple!
📁 The input files are converted into PDF by PdfItDown (https://github.com/AstraBert/PdfItDown)
📑 The PDF text is extracted using LlamaIndex readers
🦛 The text is chunked exploiting Chonkie
🧮 The chunks are embedded thanks to Sentence Transformers models
🗄️ The embeddings are loaded into a Qdrant vector database
And you're done!✅
Curious of trying it? Install it by running:
𝘱𝘪𝘱 𝘪𝘯𝘴𝘵𝘢𝘭𝘭 𝘪𝘯𝘨𝘦𝘴𝘵-𝘢𝘯𝘺𝘵𝘩𝘪𝘯𝘨
And you can start using it in your python scripts!🐍
Don't forget to star it on GitHub and let me know if you have any feedback! ➡️ https://github.com/AstraBert/ingest-anything
Hey @T-2000 , you're absolutely right! I'm in the process of making the application online so for now the repo got a bit messy, tomorrow it will be clean and ready to be spinned up also locally: sorry for the incovenient!
That's why I decided to build 𝐑𝐞𝐬𝐮𝐦𝐞 𝐌𝐚𝐭𝐜𝐡𝐞𝐫 (https://github.com/AstraBert/resume-matcher), a fully open-source application that scans your resume and searches the web for jobs that match with it!🎉
The workflow is very simple:
🦙 A LlamaExtract agent parses the resume and extracts valuable data that represent your profile
🗄️The structured data are passed on to a Job Matching Agent (built with LlamaIndex😉) that uses them to build a web search query based on your resume
🌐 The web search is handled by Linkup, which finds the top matches and returns them to the Agent
🔎 The agent evaluates the match between your profile and the jobs, and then returns a final answer to you
So, are you ready to find a job suitable for you?💼 You can spin up the application completely locally and with Docker, starting from the GitHub repo ➡️ https://github.com/AstraBert/resume-matcher
Feel free to leave your feedback and let me know in the comments if you want an online version of Resume Matcher as well!✨
I used good old Canva (pro :)
The workflow behind 𝗟𝗹𝗮𝗺𝗮𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵𝗲𝗿 is simple:
💬 You submit a query
🛡️ Your query is evaluated by Llama 3 guard model, which deems it safe or unsafe
🧠 If your query is safe, it is routed to the Researcher Agent
⚙️ The Researcher Agent expands the query into three sub-queries, with which to search the web
🌐 The web is searched for each of the sub-queries
📊 The retrieved information is evaluated for relevancy against your original query
✍️ The Researcher Agent produces an essay based on the information it gathered, paying attention to referencing its sources
The agent itself is also built with easy-to-use and intuitive blocks:
🦙 LlamaIndex provides the agentic architecture and the integrations with the language models
⚡Groq makes Llama-4 available with its lightning-fast inference
🔎 Linkup allows the agent to deep-search the web and provides sourced answers
💪 FastAPI does the heavy loading with wrapping everything within an elegant API interface
⏱️ Redis is used for API rate limiting
🎨 Gradio creates a simple but powerful user interface
Special mention also to Lovable, which helped me build the first draft of the landing page for LlamaResearcher!💖
If you're curious and you want to try LlamaResearcher, you can - completely for free and without subscription - for 30 days from now ➡️ https://llamaresearcher.com
And if you're like me, and you like getting your hands in code and build stuff on your own machine, I have good news: this is all open-source, fully reproducible locally and Docker-ready🐋
Just go to the GitHub repo: https://github.com/AstraBert/llama-4-researcher and don't forget to star it, if you find it useful!⭐
As always, have fun and feel free to leave your feedback✨
Meet 𝐓𝐲𝐒𝐕𝐀 (𝗧𝘆pe𝗦cript 𝗩oice 𝗔ssistant, https://github.com/AstraBert/TySVA), your (speaking) AI companion for everyday TypeScript programming tasks!🎙️
TySVA is a skilled TypeScript expert and, to provide accurate and up-to-date responses, she leverages the following workflow:
🗣️ If you talk to her, she converts the audio into a textual prompt, and use it a starting point to answer your questions (if you send a message, she'll use directly that💬)
🧠 She can solve your questions by (deep)searching the web and/or by retrieving relevant information from a vector database containing TypeScript documentation. If the answer is simple, she can also reply directly (no tools needed!)
🛜 To ease her life, TySVA has all the tools she needs available through Model Context Protocol (MCP)
🔊 Once she's done, she returns her answer to you, along with a voice summary of what she did and what solution she found
But how does she do that? What are her components?🤨
📖 Qdrant + HuggingFace give her the documentation knowledge, providing the vector database and the embeddings
🌐 Linkup provides her with up-to-date, grounded answers, connecting her to the web
🦙 LlamaIndex makes up her brain, with the whole agentic architecture
🎤 ElevenLabs gives her ears and mouth, transcribing and producing voice inputs and outoputs
📜 Groq provides her with speech, being the LLM provider behind TySVA
🎨 Gradio+FastAPI make up her face and fibers, providing a seamless backend-to-frontend integration
If you're now curious of trying her, you can easily do that by spinning her up locally (and with Docker!🐋) from the GitHub repo ➡️ https://github.com/AstraBert/TySVA
And feel free to leave any feedback!✨