Skip to content
Home » pySpark » How to List Files and Tables in Microsoft Fabric Lakehouse Using PySpark

How to List Files and Tables in Microsoft Fabric Lakehouse Using PySpark

Rate this post

When working with Microsoft Fabric Lakehouse, you may need to check which files, folders, and tables are available in your Lakehouse. This is particularly useful when working with multiple datasets or when you want to verify whether files and tables have been created successfully.

In this tutorial, we will learn how to list files, folders, and tables available in a Microsoft Fabric Lakehouse using PySpark.

We will use a Microsoft Fabric notebook and PySpark code to explore the Lakehouse contents.

Prerequisites

Before getting started, make sure you have the following:

  • A Microsoft Fabric workspace.
  • A Lakehouse created in your Fabric workspace.
  • A Microsoft Fabric notebook attached to your Lakehouse.
  • Some existing files and tables in your Lakehouse for demonstration.

For this demonstration, I am using the existing Lakehouse named LH_Test, which contains sample files and tables created in my previous tutorials.

If you are new to Microsoft Fabric Lakehouse or don’t know how to attach a Lakehouse to a notebook, you can refer to my previous Microsoft Fabric tutorials.

1. How to List Tables in Microsoft Fabric Lakehouse Using PySpark

First, let’s understand how to retrieve the list of tables available in our Lakehouse.

In Microsoft Fabric, we can use the spark.catalog.listTables() method to get a list of tables available in the current Spark catalog.

Let’s look at the following code:

# List Tables in the Lakehouse

tables = spark.catalog.listTables()

for table in tables:
    print(table.name)

Explanation:

  • spark.catalog.listTables() retrieves the list of tables available in the current Spark catalog.
  • for table in tables iterates through each table returned by the method.
  • table.name returns the name of the table.
  • print(table.name) displays the table name in the notebook output.

When you execute this code, you will see the names of the available tables in your Lakehouse.

Display Additional Table Information

Apart from the table name, you can also retrieve additional information about each table, such as its table type and database.

Let’s modify our previous code:

# List Tables in the Lakehouse

tables = spark.catalog.listTables()

for table in tables:
    print("Table Name:", table.name)
    print("Table Type:", table.tableType)
    print("Database:", table.database)
    print()

Explanation:

In this example, we are using the same spark.catalog.listTables() method, but instead of displaying only the table name, we are retrieving three properties:

  • table.name: Returns the name of the table.
  • table.tableType: Returns the type of the table, such as MANAGED or EXTERNAL.
  • table.database: Returns the database or schema associated with the table.

The print() statement at the end adds a blank line between each table’s information, making the output easier to read.

When you execute this code, you will see the table name, table type, and database for each available table.



2. How to List Folders in Microsoft Fabric Lakehouse Using PySpark

Now let’s move to the Files section of our Lakehouse.

In Microsoft Fabric Lakehouse, we can use mssparkutils.fs.ls() to list the contents of a specified file system path.

First, let’s see how to list all folders and files available directly inside the Files section.

# List Folders in the Lakehouse

files = mssparkutils.fs.ls("Files/")

for file in files:
    print(file.name)

Explanation:

  • mssparkutils.fs.ls("Files/") retrieves the list of items available inside the Files section of the attached Lakehouse.
  • for file in files iterates through each returned item.
  • file.name returns the name of the folder or file.
  • print(file.name) displays the name in the notebook output.

When you execute this code, you will see the folders and files available directly under the Files section of your Lakehouse.

This is useful when you have multiple folders and want to quickly explore the available data without opening the Lakehouse manually.

3. How to List Files from a Specific Folder in Microsoft Fabric Lakehouse

What if you want to list files from a particular folder instead of retrieving all the contents from the Files section?

For example, in my Lakehouse, I have a folder named SampleDataCSV, which contains CSV files.

To retrieve the files available inside this specific folder, we just need to provide the folder path in the mssparkutils.fs.ls() method.

Let’s look at the code:

# List Files in the Lakehouse

files = mssparkutils.fs.ls("Files/SampleDataCSV")

for file in files:
    print(file.name)




Explanation:

  • mssparkutils.fs.ls("Files/SampleDataCSV") retrieves the list of items available inside the SampleDataCSV folder.
  • for file in files iterates through each item returned by the method.
  • file.name returns the name of each file.
  • print(file.name) displays the file names in the notebook output.

When you execute this code, you will see all the files available inside the SampleDataCSV folder.

Similarly, you can replace SampleDataCSV with any other folder name available in your Lakehouse to retrieve its contents.

Conclusion

In this tutorial, we learned how to list tables, folders, and files available in a Microsoft Fabric Lakehouse using PySpark.

Here is a quick recap of the methods and properties covered:

Method / Property Description
spark.catalog.listTables() Retrieves tables from the current Spark catalog.
table.name Returns the table name.
table.tableType Returns the table type.
table.database Returns the database or schema name.
mssparkutils.fs.ls() Lists files and folders at a specified path.
file.name Returns the name of a file or folder.

These methods are useful when exploring Lakehouse contents, checking available datasets, and managing files and tables during your PySpark development.

Video Reference:

 

Loading

Leave a Reply

Discover more from Power BI Docs

Subscribe now to keep reading and get access to the full archive.

Continue reading