已开启
回合上游补丁,数量:5个 #59
zhangxingrong创建于 2024年8月16日
回合上游补丁,数量:5个 #59
已开启
zhangxingrong创建于 2024年8月16日
从refs/pull/59/head合入到master
共 6 个文件变更+453-1
@@ -0,0 +1,80 @@
1+From 6468f96ea42f6efe42033507c4e26600b751bfcc Mon Sep 17 00:00:00 2001
2+From: Emil Ejbyfeldt <eejbyfeldt@liveintent.com>
3+Date: Thu, 5 Oct 2023 09:41:08 +0900
4+Subject: [PATCH] [SPARK-45386][SQL][3.5] Fix correctness issue with persist
5+ using StorageLevel.NONE on Dataset
6+ 
7+### What changes were proposed in this pull request?
8+Support for InMememoryTableScanExec in AQE was added in #39624, but this patch contained a bug when a Dataset is persisted using `StorageLevel.NONE`. Before that patch a query like:
9+```
10+import org.apache.spark.storage.StorageLevel
11+spark.createDataset(Seq(1, 2)).persist(StorageLevel.NONE).count()
12+```
13+would correctly return 2. But after that patch it incorrectly returns 0. This is because AQE incorrectly determines based on the runtime statistics that are collected here:
14+https://github.com/apache/spark/blob/eac5a8c7e6da94bb27e926fc9a681aed6582f7d3/sql/core/src/main/scala/org/apache/spark/sql/execution/columnar/InMemoryRelation.scala#L294
15+that the input is empty. The problem is that the action that should make sure the statistics are collected here
16+https://github.com/apache/spark/blob/eac5a8c7e6da94bb27e926fc9a681aed6582f7d3/sql/core/src/main/scala/org/apache/spark/sql/execution/adaptive/QueryStageExec.scala#L285-L291
17+never use the iterator and when we have `StorageLevel.NONE` the persisting will also not use the iterator and we will not gather the correct statistics.
18+ 
19+The proposed fix in the patch just make calling persist with StorageLevel.NONE a no-op. Changing the action since it always "emptied" the iterator would also work but seems like that would be unnecessary work in a lot of normal circumstances.
20+ 
21+### Why are the changes needed?
22+The current code has a correctness issue.
23+ 
24+### Does this PR introduce _any_ user-facing change?
25+Yes, fixes the correctness issue.
26+ 
27+### How was this patch tested?
28+New and existing unit tests.
29+ 
30+### Was this patch authored or co-authored using generative AI tooling?
31+No
32+ 
33+Closes #43213 from eejbyfeldt/SPARK-45386-branch-3.5.
34+ 
35+Authored-by: Emil Ejbyfeldt <eejbyfeldt@liveintent.com>
36+Signed-off-by: Hyukjin Kwon <gurwls223@apache.org>
37+---
38+ .../scala/org/apache/spark/sql/execution/CacheManager.scala | 4 +++-
39+ .../src/test/scala/org/apache/spark/sql/DatasetSuite.scala | 6 ++++++
40+ 2 files changed, 9 insertions(+), 1 deletion(-)
41+ 
42+diff --git a/sql/core/src/main/scala/org/apache/spark/sql/execution/CacheManager.scala b/sql/core/src/main/scala/org/apache/spark/sql/execution/CacheManager.scala
43+index 064819275e004..e906c74f8a5ee 100644
44+--- a/sql/core/src/main/scala/org/apache/spark/sql/execution/CacheManager.scala
45++++ b/sql/core/src/main/scala/org/apache/spark/sql/execution/CacheManager.scala
46+@@ -113,7 +113,9 @@ class CacheManager extends Logging with AdaptiveSparkPlanHelper {
47+ planToCache: LogicalPlan,
48+ tableName: Option[String],
49+ storageLevel: StorageLevel): Unit = {
50+- if (lookupCachedData(planToCache).nonEmpty) {
51++ if (storageLevel == StorageLevel.NONE) {
52++ // Do nothing for StorageLevel.NONE since it will not actually cache any data.
53++ } else if (lookupCachedData(planToCache).nonEmpty) {
54+ logWarning("Asked to cache already cached data.")
55+ } else {
56+ val sessionWithConfigsOff = getOrCloneSessionWithConfigsOff(spark)
57+diff --git a/sql/core/src/test/scala/org/apache/spark/sql/DatasetSuite.scala b/sql/core/src/test/scala/org/apache/spark/sql/DatasetSuite.scala
58+index c967540541a5c..6d9c43f866a0c 100644
59+--- a/sql/core/src/test/scala/org/apache/spark/sql/DatasetSuite.scala
60++++ b/sql/core/src/test/scala/org/apache/spark/sql/DatasetSuite.scala
61+@@ -45,6 +45,7 @@ import org.apache.spark.sql.functions._
62+ import org.apache.spark.sql.internal.SQLConf
63+ import org.apache.spark.sql.test.SharedSparkSession
64+ import org.apache.spark.sql.types._
65++import org.apache.spark.storage.StorageLevel
66+
67+ case class TestDataPoint(x: Int, y: Double, s: String, t: TestDataPoint2)
68+ case class TestDataPoint2(x: Int, s: String)
69+@@ -2535,6 +2536,11 @@ class DatasetSuite extends QueryTest
70+
71+ checkDataset(ds.filter(f(col("_1"))), Tuple1(ValueClass(2)))
72+ }
73++
74++ test("SPARK-45386: persist with StorageLevel.NONE should give correct count") {
75++ val ds = Seq(1, 2).toDS().persist(StorageLevel.NONE)
76++ assert(ds.count() == 2)
77++ }
78+ }
79+
80+ class DatasetLargeResultCollectingSuite extends QueryTest
@@ -0,0 +1,56 @@
1+From 24f88b319c88bfe55e8b2b683193a85842bdad88 Mon Sep 17 00:00:00 2001
2+From: yorksity <yorksity@outlook.com>
3+Date: Tue, 10 Oct 2023 14:36:23 +0800
4+Subject: [PATCH] [SPARK-45205][SQL] CommandResultExec to override iterator
5+ methods to avoid triggering multiple jobs
6+MIME-Version: 1.0
7+Content-Type: text/plain; charset=UTF-8
8+Content-Transfer-Encoding: 8bit
9+ 
10+### What changes were proposed in this pull request?
11+ 
12+After SPARK-35378 was changed, the execution of statements such as ‘show parititions test' became slower. The change point is that the execution process changes from ExecutedCommandEnec to CommandResultExec, but ExecutedCommandExec originally implemented the following method
13+ 
14+override def executeToIterator(): Iterator[InternalRow] = sideEffectResult.iterator
15+ 
16+CommandResultExec is not rewritten, so when the hasNext method is executed, a job process is created, resulting in increased time-consuming
17+ 
18+### Why are the changes needed?
19+ 
20+Improve performance when show partitions/tables.
21+ 
22+### Does this PR introduce _any_ user-facing change?
23+ 
24+No
25+ 
26+### How was this patch tested?
27+ 
28+Existing tests should cover this.
29+ 
30+### Was this patch authored or co-authored using generative AI tooling?
31+ 
32+No
33+ 
34+Closes #43270 from yorksity/SPARK-45205.
35+ 
36+Authored-by: yorksity <yorksity@outlook.com>
37+Signed-off-by: Wenchen Fan <wenchen@databricks.com>
38+(cherry picked from commit c9c99222e828d556552694dfb48c75bf0703a2c4)
39+Signed-off-by: Wenchen Fan <wenchen@databricks.com>
40+---
41+ .../org/apache/spark/sql/execution/CommandResultExec.scala | 2 ++
42+ 1 file changed, 2 insertions(+)
43+ 
44+diff --git a/sql/core/src/main/scala/org/apache/spark/sql/execution/CommandResultExec.scala b/sql/core/src/main/scala/org/apache/spark/sql/execution/CommandResultExec.scala
45+index 5f38278d2dc67..45e3e41ab053d 100644
46+--- a/sql/core/src/main/scala/org/apache/spark/sql/execution/CommandResultExec.scala
47++++ b/sql/core/src/main/scala/org/apache/spark/sql/execution/CommandResultExec.scala
48+@@ -81,6 +81,8 @@ case class CommandResultExec(
49+ unsafeRows
50+ }
51+
52++ override def executeToIterator(): Iterator[InternalRow] = unsafeRows.iterator
53++
54+ override def executeTake(limit: Int): Array[InternalRow] = {
55+ val taken = unsafeRows.take(limit)
56+ longMetric("numOutputRows").add(taken.size)
@@ -0,0 +1,98 @@
1+From 81a7f8f184cd597208fcad72130354288a0c9f79 Mon Sep 17 00:00:00 2001
2+From: liangyongyuan <liangyongyuan@xiaomi.com>
3+Date: Tue, 10 Oct 2023 14:40:33 +0800
4+Subject: [PATCH] [SPARK-45449][SQL] Cache Invalidation Issue with JDBC Table
5+ 
6+### What changes were proposed in this pull request?
7+Add an equals method to `JDBCOptions` that considers two instances equal if their `JDBCOptions.parameters` are the same.
8+ 
9+### Why are the changes needed?
10+We have identified a cache invalidation issue when caching JDBC tables in Spark SQL. The cached table is unexpectedly invalidated when queried, leading to a re-read from the JDBC table instead of retrieving data from the cache.
11+Example SQL:
12+ 
13+```
14+CACHE TABLE cache_t SELECT * FROM mysql.test.test1;
15+SELECT * FROM cache_t;
16+```
17+Expected Behavior:
18+The expectation is that querying the cached table (cache_t) should retrieve the result from the cache without re-evaluating the execution plan.
19+ 
20+Actual Behavior:
21+However, the cache is invalidated, and the content is re-read from the JDBC table.
22+ 
23+Root Cause:
24+The issue lies in the `CacheData` class, where the comparison involves `JDBCTable`. The `JDBCTable` is a case class:
25+ 
26+`case class JDBCTable(ident: Identifier, schema: StructType, jdbcOptions: JDBCOptions)`
27+The comparison of non-case class components, such as `jdbcOptions`, involves pointer comparison. This leads to unnecessary cache invalidation.
28+ 
29+### Does this PR introduce _any_ user-facing change?
30+No
31+ 
32+### How was this patch tested?
33+Add uts
34+ 
35+### Was this patch authored or co-authored using generative AI tooling?
36+No
37+ 
38+Closes #43258 from lyy-pineapple/spark-git-cache.
39+ 
40+Authored-by: liangyongyuan <liangyongyuan@xiaomi.com>
41+Signed-off-by: Wenchen Fan <wenchen@databricks.com>
42+(cherry picked from commit d073f2d3e2f67a4b612e020a583e23dc1fa63aab)
43+Signed-off-by: Wenchen Fan <wenchen@databricks.com>
44+---
45+ .../execution/datasources/jdbc/JDBCOptions.scala | 8 ++++++++
46+ .../v2/jdbc/JDBCTableCatalogSuite.scala | 15 +++++++++++++++
47+ 2 files changed, 23 insertions(+)
48+ 
49+diff --git a/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/jdbc/JDBCOptions.scala b/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/jdbc/JDBCOptions.scala
50+index 268a65b81ff68..57651684070f7 100644
51+--- a/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/jdbc/JDBCOptions.scala
52++++ b/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/jdbc/JDBCOptions.scala
53+@@ -239,6 +239,14 @@ class JDBCOptions(
54+ .get(JDBC_PREFER_TIMESTAMP_NTZ)
55+ .map(_.toBoolean)
56+ .getOrElse(SQLConf.get.timestampType == TimestampNTZType)
57++
58++ override def hashCode: Int = this.parameters.hashCode()
59++
60++ override def equals(other: Any): Boolean = other match {
61++ case otherOption: JDBCOptions =>
62++ otherOption.parameters.equals(this.parameters)
63++ case _ => false
64++ }
65+ }
66+
67+ class JdbcOptionsInWrite(
68+diff --git a/sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/v2/jdbc/JDBCTableCatalogSuite.scala b/sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/v2/jdbc/JDBCTableCatalogSuite.scala
69+index 6b85911dca773..eed64b873c451 100644
70+--- a/sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/v2/jdbc/JDBCTableCatalogSuite.scala
71++++ b/sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/v2/jdbc/JDBCTableCatalogSuite.scala
72+@@ -26,6 +26,7 @@ import org.apache.spark.sql.{AnalysisException, QueryTest, Row}
73+ import org.apache.spark.sql.catalyst.analysis.{NoSuchNamespaceException, TableAlreadyExistsException}
74+ import org.apache.spark.sql.catalyst.parser.ParseException
75+ import org.apache.spark.sql.catalyst.util.CharVarcharUtils
76++import org.apache.spark.sql.execution.columnar.InMemoryTableScanExec
77+ import org.apache.spark.sql.internal.SQLConf
78+ import org.apache.spark.sql.test.SharedSparkSession
79+ import org.apache.spark.sql.types._
80+@@ -512,4 +513,18 @@ class JDBCTableCatalogSuite extends QueryTest with SharedSparkSession {
81+ assert(t.schema === replaced)
82+ }
83+ }
84++
85++ test("SPARK-45449: Cache Invalidation Issue with JDBC Table") {
86++ withTable("h2.test.cache_t") {
87++ withConnection { conn =>
88++ conn.prepareStatement(
89++ """CREATE TABLE "test"."cache_t" (id decimal(25) PRIMARY KEY NOT NULL,
90++ |name TEXT(32) NOT NULL)""".stripMargin).executeUpdate()
91++ }
92++ sql("INSERT OVERWRITE h2.test.cache_t SELECT 1 AS id, 'a' AS name")
93++ sql("CACHE TABLE t1 SELECT id, name FROM h2.test.cache_t")
94++ val plan = sql("select * from t1").queryExecution.sparkPlan
95++ assert(plan.isInstanceOf[InMemoryTableScanExec])
96++ }
97++ }
98+ }
@@ -0,0 +1,116 @@
1+From 6a5747d66e53ed0d934cdd9ca5c9bd9fde6868e6 Mon Sep 17 00:00:00 2001
2+From: Kent Yao <yao@apache.org>
3+Date: Tue, 17 Oct 2023 22:19:18 +0800
4+Subject: [PATCH] [SPARK-45568][TESTS] Fix flaky
5+ WholeStageCodegenSparkSubmitSuite
6+ 
7+### What changes were proposed in this pull request?
8+ 
9+WholeStageCodegenSparkSubmitSuite is [flaky](https://github.com/apache/spark/actions/runs/6479534195/job/17593342589) because SHUFFLE_PARTITIONS(200) creates 200 reducers for one total core and improper stop progress causes executor launcher reties. The heavy load and reties might result in timeout test failures.
10+ 
11+### Why are the changes needed?
12+ 
13+CI robustness
14+ 
15+### Does this PR introduce _any_ user-facing change?
16+ 
17+no
18+ 
19+### How was this patch tested?
20+ 
21+existing WholeStageCodegenSparkSubmitSuite
22+### Was this patch authored or co-authored using generative AI tooling?
23+ 
24+no
25+ 
26+Closes #43394 from yaooqinn/SPARK-45568.
27+ 
28+Authored-by: Kent Yao <yao@apache.org>
29+Signed-off-by: Kent Yao <yao@apache.org>
30+(cherry picked from commit f00ec39542a5f9ac75d8c24f0f04a7be703c8d7c)
31+Signed-off-by: Kent Yao <yao@apache.org>
32+---
33+ .../WholeStageCodegenSparkSubmitSuite.scala | 57 ++++++++++---------
34+ 1 file changed, 30 insertions(+), 27 deletions(-)
35+ 
36+diff --git a/sql/core/src/test/scala/org/apache/spark/sql/execution/WholeStageCodegenSparkSubmitSuite.scala b/sql/core/src/test/scala/org/apache/spark/sql/execution/WholeStageCodegenSparkSubmitSuite.scala
37+index e253de76221ad..69145d890fc19 100644
38+--- a/sql/core/src/test/scala/org/apache/spark/sql/execution/WholeStageCodegenSparkSubmitSuite.scala
39++++ b/sql/core/src/test/scala/org/apache/spark/sql/execution/WholeStageCodegenSparkSubmitSuite.scala
40+@@ -26,6 +26,7 @@ import org.apache.spark.deploy.SparkSubmitTestUtils
41+ import org.apache.spark.internal.Logging
42+ import org.apache.spark.sql.{QueryTest, Row, SparkSession}
43+ import org.apache.spark.sql.functions.{array, col, count, lit}
44++import org.apache.spark.sql.internal.SQLConf
45+ import org.apache.spark.sql.types.IntegerType
46+ import org.apache.spark.tags.ExtendedSQLTest
47+ import org.apache.spark.unsafe.Platform
48+@@ -70,39 +71,41 @@ class WholeStageCodegenSparkSubmitSuite extends SparkSubmitTestUtils
49+
50+ object WholeStageCodegenSparkSubmitSuite extends Assertions with Logging {
51+
52+- var spark: SparkSession = _
53+-
54+ def main(args: Array[String]): Unit = {
55+ TestUtils.configTestLog4j2("INFO")
56+
57+- spark = SparkSession.builder().getOrCreate()
58++ val spark = SparkSession.builder()
59++ .config(SQLConf.SHUFFLE_PARTITIONS.key, "2")
60++ .getOrCreate()
61++
62++ try {
63++ // Make sure the test is run where the driver and the executors uses different object layouts
64++ val driverArrayHeaderSize = Platform.BYTE_ARRAY_OFFSET
65++ val executorArrayHeaderSize =
66++ spark.sparkContext.range(0, 1).map(_ => Platform.BYTE_ARRAY_OFFSET).collect().head
67++ assert(driverArrayHeaderSize > executorArrayHeaderSize)
68+
69+- // Make sure the test is run where the driver and the executors uses different object layouts
70+- val driverArrayHeaderSize = Platform.BYTE_ARRAY_OFFSET
71+- val executorArrayHeaderSize =
72+- spark.sparkContext.range(0, 1).map(_ => Platform.BYTE_ARRAY_OFFSET).collect.head.toInt
73+- assert(driverArrayHeaderSize > executorArrayHeaderSize)
74++ val df = spark.range(71773).select((col("id") % lit(10)).cast(IntegerType) as "v")
75++ .groupBy(array(col("v"))).agg(count(col("*")))
76++ val plan = df.queryExecution.executedPlan
77++ assert(plan.exists(_.isInstanceOf[WholeStageCodegenExec]))
78+
79+- val df = spark.range(71773).select((col("id") % lit(10)).cast(IntegerType) as "v")
80+- .groupBy(array(col("v"))).agg(count(col("*")))
81+- val plan = df.queryExecution.executedPlan
82+- assert(plan.exists(_.isInstanceOf[WholeStageCodegenExec]))
83++ val expectedAnswer =
84++ Row(Array(0), 7178) ::
85++ Row(Array(1), 7178) ::
86++ Row(Array(2), 7178) ::
87++ Row(Array(3), 7177) ::
88++ Row(Array(4), 7177) ::
89++ Row(Array(5), 7177) ::
90++ Row(Array(6), 7177) ::
91++ Row(Array(7), 7177) ::
92++ Row(Array(8), 7177) ::
93++ Row(Array(9), 7177) :: Nil
94+
95+- val expectedAnswer =
96+- Row(Array(0), 7178) ::
97+- Row(Array(1), 7178) ::
98+- Row(Array(2), 7178) ::
99+- Row(Array(3), 7177) ::
100+- Row(Array(4), 7177) ::
101+- Row(Array(5), 7177) ::
102+- Row(Array(6), 7177) ::
103+- Row(Array(7), 7177) ::
104+- Row(Array(8), 7177) ::
105+- Row(Array(9), 7177) :: Nil
106+- val result = df.collect
107+- QueryTest.sameRows(result.toSeq, expectedAnswer) match {
108+- case Some(errMsg) => fail(errMsg)
109+- case _ =>
110++ QueryTest.checkAnswer(df, expectedAnswer)
111++ } finally {
112++ spark.stop()
113+ }
114++
115+ }
116+ }
@@ -0,0 +1,84 @@
1+From f47b63c6a62fb6f1fd894f64736847719af7a199 Mon Sep 17 00:00:00 2001
2+From: allisonwang-db <allison.wang@databricks.com>
3+Date: Fri, 20 Oct 2023 08:36:42 +0800
4+Subject: [PATCH] [SPARK-45584][SQL] Fix subquery execution failure with
5+ TakeOrderedAndProjectExec
6+ 
7+This PR fixes a bug when there are subqueries in `TakeOrderedAndProjectExec`. The executeCollect method does not wait for subqueries to finish and it can result in IllegalArgumentException when executing a simple query.
8+For example this query:
9+```
10+WITH t2 AS (
11+ SELECT * FROM t1 ORDER BY id
12+)
13+SELECT *, (SELECT COUNT(*) FROM t2) FROM t2 LIMIT 10
14+```
15+will fail with this error
16+```
17+ java.lang.IllegalArgumentException: requirement failed: Subquery subquery#242, [id=#109] has not finished
18+```
19+ 
20+To fix a bug.
21+ 
22+No
23+ 
24+New unit test
25+ 
26+No
27+ 
28+Closes #43419 from allisonwang-db/spark-45584-subquery-failure.
29+ 
30+Authored-by: allisonwang-db <allison.wang@databricks.com>
31+Signed-off-by: Wenchen Fan <wenchen@databricks.com>
32+(cherry picked from commit 8fd915ffaba1cc99813cc8d6d2a28688d7fae39b)
33+Signed-off-by: Wenchen Fan <wenchen@databricks.com>
34+---
35+ .../apache/spark/sql/execution/limit.scala | 2 +-
36+ .../org/apache/spark/sql/SubquerySuite.scala | 24 +++++++++++++++++++
37+ 2 files changed, 25 insertions(+), 1 deletion(-)
38+ 
39+diff --git a/sql/core/src/main/scala/org/apache/spark/sql/execution/limit.scala b/sql/core/src/main/scala/org/apache/spark/sql/execution/limit.scala
40+index 877f6508d963f..77135d21a26ab 100644
41+--- a/sql/core/src/main/scala/org/apache/spark/sql/execution/limit.scala
42++++ b/sql/core/src/main/scala/org/apache/spark/sql/execution/limit.scala
43+@@ -282,7 +282,7 @@ case class TakeOrderedAndProjectExec(
44+ projectList.map(_.toAttribute)
45+ }
46+
47+- override def executeCollect(): Array[InternalRow] = {
48++ override def executeCollect(): Array[InternalRow] = executeQuery {
49+ val orderingSatisfies = SortOrder.orderingSatisfies(child.outputOrdering, sortOrder)
50+ val ord = new LazilyGeneratedOrdering(sortOrder, child.output)
51+ val limited = if (orderingSatisfies) {
52+diff --git a/sql/core/src/test/scala/org/apache/spark/sql/SubquerySuite.scala b/sql/core/src/test/scala/org/apache/spark/sql/SubquerySuite.scala
53+index d235d2a15fea3..a7a0f6156cb1d 100644
54+--- a/sql/core/src/test/scala/org/apache/spark/sql/SubquerySuite.scala
55++++ b/sql/core/src/test/scala/org/apache/spark/sql/SubquerySuite.scala
56+@@ -2712,4 +2712,28 @@ class SubquerySuite extends QueryTest
57+ expected)
58+ }
59+ }
60++
61++ test("SPARK-45584: subquery execution should not fail with ORDER BY and LIMIT") {
62++ withTable("t1") {
63++ sql(
64++ """
65++ |CREATE TABLE t1 USING PARQUET
66++ |AS SELECT * FROM VALUES
67++ |(1, "a"),
68++ |(2, "a"),
69++ |(3, "a") t(id, value)
70++ |""".stripMargin)
71++ val df = sql(
72++ """
73++ |WITH t2 AS (
74++ | SELECT * FROM t1 ORDER BY id
75++ |)
76++ |SELECT *, (SELECT COUNT(*) FROM t2) FROM t2 LIMIT 10
77++ |""".stripMargin)
78++ // This should not fail with IllegalArgumentException.
79++ checkAnswer(
80++ df,
81++ Row(1, "a", 3) :: Row(2, "a", 3) :: Row(3, "a", 3) :: Nil)
82++ }
83++ }
84+ }
@@ -4,7 +4,7 @@
4Summary: A unified analytics engine for large-scale data processing.4Summary: A unified analytics engine for large-scale data processing.
5Name: spark5Name: spark
6Version: 3.5.06Version: 3.5.0
7-Release: 47+Release: 5
8License: Apache 2.08License: Apache 2.0
9URL: http://spark.apache.org/9URL: http://spark.apache.org/
10Source0: https://github.com/apache/spark/archive/v%{version}.tar.gz10Source0: https://github.com/apache/spark/archive/v%{version}.tar.gz
@@ -17,6 +17,12 @@ Source6: https://github.com/grpc/grpc-java/archive/refs/tags/v1.56.0.tar.gz
17Patch0001: 0001-change-mvn-scalafmt.patch17Patch0001: 0001-change-mvn-scalafmt.patch
18Patch0002: 0002-Upgrade-os-maven-plugin-to-1.7.1.patch18Patch0002: 0002-Upgrade-os-maven-plugin-to-1.7.1.patch
19 19 
20+Patch0003: 0003-Fix-correctness-issue-with-persist-using-StorageLevel.NONE-on-Dataset.patch
21+Patch0004: 0004-CommandResultExec-to-override-iterator-methods-to-avoid-triggering-multiple-jobs.patch
22+Patch0005: 0005-Cache-Invalidation-Issue-with-JDBC-Table.patch
23+Patch0006: 0006-Fix-flaky-WholeStageCodegenSparkSubmitSuite.patch
24+Patch0007: 0007-Fix-subquery-execution-failure-with-TakeOrderedAndProjectExec.patch
25+ 
20%ifarch riscv6426%ifarch riscv64
21BuildRequires: protobuf-devel protobuf-compiler27BuildRequires: protobuf-devel protobuf-compiler
22BuildRequires: autoconf automake libtool pkgconfig zlib-devel libstdc++-static gcc-c++28BuildRequires: autoconf automake libtool pkgconfig zlib-devel libstdc++-static gcc-c++
@@ -76,6 +82,11 @@ popd
76 82 
77%patch0001 -p183%patch0001 -p1
78%patch0002 -p184%patch0002 -p1
85+%patch0003 -p1
86+%patch0004 -p1
87+%patch0005 -p1
88+%patch0006 -p1
89+%patch0007 -p1
79 90 
80%ifarch riscv6491%ifarch riscv64
81sed -i -e 's/protoVersion = "3.23.4/protoVersion = "'${PROTOC_VERSION}/'' project/SparkBuild.scala92sed -i -e 's/protoVersion = "3.23.4/protoVersion = "'${PROTOC_VERSION}/'' project/SparkBuild.scala
@@ -97,6 +108,13 @@ cp -rf ../%{name}-%{version} %{buildroot}/opt/apache-%{name}-%{version}
97 108 
98 109 
99%changelog110%changelog
111+* Fri Aug 16 2024 zhangxingrong <zhangxingrong@uniontech.cn> - 3.5.0-5
112+- Fix correctness issue with persist using StorageLevel.NONE on Dataset
113+- CommandResultExec to override iterator methods to avoid triggering multiple jobs
114+- Cache Invalidation Issue with JDBC Table
115+- Fix flaky WholeStageCodegenSparkSubmitSuite
116+- Fix subquery execution failure with TakeOrderedAndProjectExec
117+ 
100* Mon Jul 1 2024 Dingli Zhang <dingli@iscas.ac.cn> - 3.5.0-4118* Mon Jul 1 2024 Dingli Zhang <dingli@iscas.ac.cn> - 3.5.0-4
101- Add riscv64 to ExclusiveArch119- Add riscv64 to ExclusiveArch
102- Fix build on riscv64120- Fix build on riscv64