
Explanation:
spark.sql("<query>") executes a SQL query and returns a PySpark DataFrame, which can then be tested and transformed with Python. spark.table is appropriate when loading a table by name, but the requirement specifically describes executing an existing query.
A data analyst has a SQL query against a Delta table, and the data engineering team wants to execute that query and work with its results in PySpark. Which operation should they use?
A
SELECT * FROM sales
B
spark.delta.table
C
spark.sql
D
spark.table
No comments yet.