{ "cells": [ { "cell_type": "markdown", "id": "20deed05", "metadata": {}, "source": [ "# Unstructured File Loader\n", "This notebook covers how to use Unstructured to load files of many types. Unstructured currently supports loading of text files, powerpoints, html, pdfs, images, and more." ] }, { "cell_type": "code", "execution_count": 1, "id": "2886982e", "metadata": {}, "outputs": [], "source": [ "# # Install package\n", "# !pip install unstructured" ] }, { "cell_type": "code", "execution_count": 2, "id": "54d62efd", "metadata": {}, "outputs": [], "source": [ "# # Install other dependencies\n", "# # https://github.com/Unstructured-IO/unstructured/blob/main/docs/source/installing.rst\n", "# !brew install libmagic" ] }, { "cell_type": "code", "execution_count": 3, "id": "af6a64f5", "metadata": {}, "outputs": [], "source": [ "# import nltk\n", "# nltk.download('punkt')" ] }, { "cell_type": "code", "execution_count": 4, "id": "79d3e549", "metadata": {}, "outputs": [], "source": [ "from langchain.document_loaders import UnstructuredFileLoader" ] }, { "cell_type": "code", "execution_count": 5, "id": "2593d1dc", "metadata": {}, "outputs": [], "source": [ "loader = UnstructuredFileLoader(\"../../state_of_the_union.txt\")" ] }, { "cell_type": "code", "execution_count": 6, "id": "fe34e941", "metadata": {}, "outputs": [], "source": [ "docs = loader.load()" ] }, { "cell_type": "code", "execution_count": 7, "id": "ee449788", "metadata": {}, "outputs": [ { "data": { "text/plain": [ "'Madam Speaker, Madam Vice President, our First Lady and Second Gentleman. Members of Congress and the Cabinet. Justices of the Supreme Court. My fellow Americans.\\n\\nLast year COVID-19 kept us apart. This year we are finally together again.\\n\\nTonight, we meet as Democrats Republicans and Independents. But most importantly as Americans.\\n\\nWith a duty to one another to the American people to the Constit'" ] }, "execution_count": 7, "metadata": {}, "output_type": "execute_result" } ], "source": [ "docs[0].page_content[:400]" ] }, { "cell_type": "markdown", "id": "7874d01d", "metadata": {}, "source": [ "## Retain Elements\n", "\n", "Under the hood, Unstructured creates different \"elements\" for different chunks of text. By default we combine those together, but you can easily keep that separation by specifying `mode=\"elements\"`." ] }, { "cell_type": "code", "execution_count": 8, "id": "ff5b616d", "metadata": {}, "outputs": [], "source": [ "loader = UnstructuredFileLoader(\"../../state_of_the_union.txt\", mode=\"elements\")" ] }, { "cell_type": "code", "execution_count": 9, "id": "feca3b6c", "metadata": {}, "outputs": [], "source": [ "docs = loader.load()" ] }, { "cell_type": "code", "execution_count": 12, "id": "fec5bbac", "metadata": {}, "outputs": [ { "data": { "text/plain": [ "[Document(page_content='Madam Speaker, Madam Vice President, our First Lady and Second Gentleman. Members of Congress and the Cabinet. Justices of the Supreme Court. My fellow Americans.', lookup_str='', metadata={'source': '../../state_of_the_union.txt'}, lookup_index=0),\n", " Document(page_content='Last year COVID-19 kept us apart. This year we are finally together again.', lookup_str='', metadata={'source': '../../state_of_the_union.txt'}, lookup_index=0),\n", " Document(page_content='Tonight, we meet as Democrats Republicans and Independents. But most importantly as Americans.', lookup_str='', metadata={'source': '../../state_of_the_union.txt'}, lookup_index=0),\n", " Document(page_content='With a duty to one another to the American people to the Constitution.', lookup_str='', metadata={'source': '../../state_of_the_union.txt'}, lookup_index=0),\n", " Document(page_content='And with an unwavering resolve that freedom will always triumph over tyranny.', lookup_str='', metadata={'source': '../../state_of_the_union.txt'}, lookup_index=0)]" ] }, "execution_count": 12, "metadata": {}, "output_type": "execute_result" } ], "source": [ "docs[:5]" ] }, { "cell_type": "code", "execution_count": null, "id": "8ca8a648", "metadata": {}, "outputs": [], "source": [] } ], "metadata": { "kernelspec": { "display_name": "Python 3 (ipykernel)", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.9.1" } }, "nbformat": 4, "nbformat_minor": 5 }